DataSeeds.AI
Rights-cleared image and video training data, produced to spec by a managed crowd of photographers and on-the-ground collectors. Part of Zedge.
What you will find here
- DSD sample dataset: 7,770 peer-ranked photographs with scene descriptions, object labels, segmentation masks and full EXIF. Apache 2.0.
- Baselines: LLaVA-OneVision and BLIP2 fine-tuned on DSD, with the training code on GitHub.
- Paper: Peer-Ranked Precision, on whether peer-ranked, aesthetically scored data fine-tunes better than scraped web data.
- Coming: an open egocentric video sample for Physical AI and robotics.
How the data is made
Every image and clip is captured by a known contributor under a consent agreement, ranked by the community, and documented for GDPR, CCPA and BIPA. Provenance travels with the file.
Need more than the sample?
Larger collections, custom production and annotation: dataseeds.ai or sales@dataseeds.ai.