Skip to content

Image data

Image Datasets for Computer Vision: How AI Teams Source Real-World Images

Quick answer

Image datasets for computer vision come from four places: public benchmarks such as ImageNet, COCO and Open Images; web-scraped collections; licensed photo archives that companies built while doing their work, such as claims, inspection and quality-control photos; and newly commissioned capture. Benchmarks suit prototyping and comparison, but their terms often do not clearly cover commercial training. Archives add real conditions and outcome labels, while commissioned collection adds control over scene, sensor and releases. Many production teams combine the last two.

By SourceX Editorial · Updated

Public, scraped, licensed or commissioned: four routes to vision data

The route decides three things before any pixel is reviewed: whether labels exist, who can grant training rights, and how closely the images match your deployment cameras.

RouteWhat you getLabelsRights positionBreaks whenBest fit
Public benchmarks and open datasetsCurated images, published splitsIncluded, in the benchmark's own taxonomyPer-dataset terms, often research-only or mixed per imageDomain, sensor or classes differ; commercial usePrototyping, comparison with published results
Web-scraped collectionsVolume and visual varietyAlt text, filenames or noneCopyright in every photo, EU opt-outs, personal data in pixels and headersYou need a documented chain of titleResearch; hard to defend for commercial training
Licensed operational archivesPhotos taken during real work and linked to recordsOutcome fields in the record; boxes or masks added laterLicensed by the holder, within its customer and staff termsCapture was inconsistent or the outcome field is noisyDomain fine-tuning, real-condition evaluation, rare real failures
Commissioned collectionNew images captured to your protocolLabeled to your guidelineAssigned or licensed under your contract, with releases you draftedYou need history, outcomes or long-tail variationMissing classes, new products or sensors, consented images of people

Synthetic and rendered images are a fifth input, usually for augmentation; the licensed vs synthetic vs scraped data guide compares them.

What public benchmarks can and cannot do for a commercial model

Public benchmarks are built so results can be compared, not so a company can ship a model trained on them, and three gaps appear when a team tries.

Licenses that do not reach every image. A 2021 study of six commonly used public image datasets found potential license-violation risk in five of them if used to build commercial AI software, partly because one dataset can combine sources under different licenses [1]. SeeFar, a geospatial collection, applies CC BY-NC 4.0 to its WorldStrat imagery and CC BY-SA 3.0 IGO to its Sentinel imagery [2]. COCO-format files carry a top-level licenses section beside images and annotations [3], so check which license applies to each image rather than trusting a dataset-level badge.

Unchecked labels. An audit of ten widely used test sets estimated label errors in at least 6% of the ImageNet validation set [4]. Label scope requires checking the files, not just documentation: for example, Ultralytics documentation for SKU-110K clarifies it provides single-class bounding boxes, with no per-SKU category labels [5], agreeing with a 2020 paper that concluded the dataset cannot be used for product recognition [6].

Conditions that are not yours. A review of manufacturing defect benchmarks notes that DAGM is synthetic, holds at most one defect per image and is statistically biased relative to real production data [7]. Benchmarks rarely match your lens, lighting, background or class balance; see camera, lens and lighting diversity.

When public data falls short, teams turn to the web: a National College of Ireland thesis on car damage detection reports that the public dataset was not enough, so the author scraped Google images [8].

Scraping solves volume, but every image arrives with its own copyright, opt-out status and often personal data, and the scraper documents none of it.

  • Copyright. The U.S. Copyright Office's Part 3 report on generative AI training, a pre-publication version from May 2025, concludes that copying works into training datasets may be prima facie infringement unless an exception such as fair use applies; the report is advisory and does not bind courts [9].
  • EU opt-outs. Providers placing general-purpose AI models on the EU market must keep a copyright policy that identifies and complies with rights reservations under Article 4(3) of the DSM Directive, and must publish a summary of their training content; these obligations applied from 2 August 2025 [10].
  • Personal data in headers. An audit of DataComp's CommonPool, a web-scraped image-text pool, found Exif tags carrying timestamps, geolocation and details about individuals, many disclosing full names; the dataset's download tool extracts Exif for every sample [11].
  • Faces. Illinois' BIPA covers scans of face geometry (not photographs themselves), requires written notice and a written release before collection, and allows $1,000 per negligent and $5,000 per intentional or reckless violation; since a 2024 amendment, collecting the same biometric from the same person by the same method more than once counts as a single violation [12]. The biometric data guide adds the Texas and Washington rules.
  • Remedies that reach the model. The FTC's 2021 final order against photo-app developer Everalbum required deleting the models and algorithms built from users' photos and videos [13]. Image diffusion models can also regenerate individual training images, including photos of people and trademarked logos [14].

SourceX does not source scraped web content or generic photos.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Operational photo archives: real outcomes from systems of record

Operational archives are historical photo archives that companies keep to run their business, stored beside the record of what happened. The link is the asset: the photo shows the condition, and the record shows what someone decided about it.

Source systemTypical imagesLabel the linked record supplies
Auto and property claims platformsFirst notice of loss (FNOL) and adjuster photosDamage findings, repair operations, estimate line items, total-loss decisions
Inspection and field-service appsAsset photos tied to a checklist item or work orderPass or fail findings, defect codes, technician notes
Automated optical inspection (AOI) stationsFixed-camera images of parts and boardsReject reasons and dispositions from the quality system
Drone and aerial inspection programsPoles, towers, roofs, bridgesInspection findings and maintenance tickets
Retail shelf audits and catalog studiosShelf photos and packshotsPlanogram compliance results, SKU attributes
Construction daily logs and punch listsProgress and QA photosIssue type, location, closure status

Industry research bodies sometimes publish this kind of real-world data: the Electric Power Research Institute (EPRI) released about 30,000 drone images of overhead distribution infrastructure, anonymized and with Exif metadata removed [15].

Archives have predictable weaknesses: mixed devices and angles, free-text labels that need normalizing, heavy imbalance toward normal cases, bursts of near-duplicate shots, and bystanders, plates or addresses in frame. Anonymize selectively: in a CVPR 2023 workshop study on detection datasets, traditional blurring or masking hurt training noticeably, especially for whole bodies, while realistic replacement of faces kept the drop minimal [16] (face and plate blurring evidence). Ask who took each photo, since staff, contractors and customers can hold different rights.

SourceX sources operational datasets from US companies on request and manages the commercial process, including the license and ongoing purchases; buyers describe the data, not the businesses, and the supplying company approves every release. Categories such as images and inspection photos are kinds of data SourceX sources, not inventory under contract, and a request does not guarantee a match. Photos are often worth more with the inspection reports that record their findings. Describe the photos and linked records you need, or read how licensing proprietary data works.

Commissioned collection: control over scene, sensor and release

Commissioned collection buys control: you set devices, lenses, lighting, viewpoints, backgrounds, class quotas and release language before the first frame. It is usually the only route to real photos of products not yet on the market or of defects you can stage safely.

It cannot buy history. Staged scenes lack the long-tail variation and outcomes, such as a repair cost or reject decision, that make archive photos valuable for evaluation. With image data collection services, lead time and per-image cost grow with the sites, people and conditions required.

Three contract points decide whether a commissioned set is usable:

  1. Releases written for AI training. As market practice, Adobe Stock requires a model release whenever a person is recognizable, including by tattoos, clothing or surroundings [17], and pocstock's policy asks for consent to secondary use including AI and machine-learning training [18]. Older releases may not mention training; see model and property releases for AI training images.
  2. Ownership. Choose between IP assignment and a license in the collection contract.
  3. A written protocol. Put capture parameters, label guidelines and acceptance tests in the statement of work for custom collection.

Many teams license an archive for outcomes and variation, then commission capture for missing classes, sensors or consented people; see the custom collection versus licensing comparison.

Match the label to the vision task

Buy the least expensive label that still teaches what the model must learn; the task sets the annotation type and the file format.

TaskLabel to buyCommon containerCheck before paying
ClassificationOne or more class labels per imageCSV or folder per classTaxonomy matches yours; multi-label allowed
Object detectionBoxes with a classCOCO JSON or YOLO text filesBox tightness rules; occlusion and truncation flags
Instance segmentationPolygons or masksCOCO polygons or RLE masks [3]Boundary tolerance; iscrowd handling
Keypoints and poseNamed points per objectCOCO person_keypoints files [3]Visibility flags; skeleton definition
Anomaly detectionNormal-only training images; masks on test defectsImage folders plus mask PNGsTest defects come from real production
VLM fine-tuning and captioningDomain text per imageJSONL or WebDataset tar shards [19]Text describes the image, not the file name

A chest-radiograph study treats bounding boxes as noisier alternatives to contour annotations [20], so a detector bought with boxes may need masks later. Operational archives add another label source, the outcome field: domain captions from work records shows how adjuster comments and inspection notes become image text, and pre-labeled versus raw images helps decide whether to buy labels or add them.

What each delivered image record should carry

Each delivered image should carry capture metadata, a stated Exif policy, rights fields and label provenance, so reviewers can audit without opening pixels. For high-risk systems under the EU AI Act, Article 10(2) expects data governance covering the origin of data, annotation, labelling and cleaning, bias examination and data gaps [21]. Regulation (EU) 2026/1744 moved high-risk application dates to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems [22].

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "image_id": "img-000418-03",
  "file": "shard-000041.tar/img-000418-03.jpg",
  "source_system": "auto_claims_photo_store",
  "linked_record": {"record_type": "claim", "record_id": "CLM-000418", "split_group": "CLM-000418"},
  "capture": {"device_class": "smartphone", "width_px": 4032, "height_px": 3024,
              "format": "jpeg", "capture_month": "2024-06",
              "viewpoint": "front_left_45", "setting": "outdoor_daylight"},
  "exif_policy": {"gps": "removed", "device_serial": "removed", "owner_name": "removed",
                  "datetime_original": "truncated_to_month", "orientation": "kept"},
  "privacy": {"faces_detected": 0, "plates_detected": 1, "plate_action": "blurred",
              "deid_method_ref": "DEID-IMG-03"},
  "rights": {"photographer_role": "staff_adjuster", "copyright_holder": "supplier",
             "recognizable_people": false, "third_party_marks": ["vehicle_badge"],
             "license_ref": "LIC-0021", "permitted_uses": ["model_training", "internal_evaluation"]},
  "labels": [
    {"type": "bbox_xywh_px", "class": "dent", "value": [812, 1404, 610, 388],
     "source": "annotation_vendor", "guideline_version": "2.1"},
    {"type": "record_outcome", "field": "repair_operation", "value": "replace_front_bumper",
     "source": "estimate_line_item"}
  ]
}

split_group keeps every photo of one claim in the same split, so near-duplicate angles cannot leak from training into test (duplication and contamination checks). exif_policy keeps orientation, which controls how the image renders, and drops GPS and serials (what to strip and keep in Exif). Ask for a datasheet too; NeurIPS 2026 requires Croissant-RAI responsible-AI metadata for its Evaluations and Datasets Track [23]. The delivery formats guide covers shards and transfer.

Start here: image data guides by domain and concern

Find your domain in the table; cross-cutting guides on volume, resolution, annotation and licensing follow it.

Medical images need extra care: DICOM states that its confidentiality profiles do not guarantee removal of all identifying information [24], and for health records SourceX requires HIPAA de-identification (Safe Harbor or Expert Determination) before anything is considered for a license.

Cross-cutting guides: how many images a model needs, resolution and compression, boxes, polygons, masks or keypoints, annotation guidelines, unlabeled corpora for pre-training and image licensing terms. Document images belong to the document AI hub, moving images to the video hub, and the AI data buyer's guide maps every cluster.

Image sourcing mistakes that surface after training

These mistakes pass a sample review and surface once a model ships.

  1. Reading the dataset badge, not per-image licenses. Mixed terms carry into commercial weights.
  2. Splitting by image instead of by object, site or claim. Near-duplicates inflate test scores.
  3. Stripping every header field. Orientation, capture month and device class support rendering and stratified evaluation.
  4. Buying volume of easy normal images. Set per-class minimums for the rare failures that matter.
  5. Testing only on benchmarks. Hold out images from your own deployment cameras.

Looking for real-world images for computer vision?

Describe the task, classes, capture conditions, linked records, volume and allowed uses on the SourceX buyer page. SourceX looks for US companies that hold matching photo archives, checks the data and the supplier's licensing permissions, agrees a license, and coordinates delivery and payment; nothing is contracted until a supplier agrees. Request real-world image data through SourceX.

Guides in this section

Sources

  1. arXiv:2111.02374, "Can I use this publicly available dataset to build commercial AI software?" (2021). https://arxiv.org/abs/2111.02374v4
  2. Registry of Open Data on AWS, "SeeFar". https://registry.opendata.aws/seefar/
  3. CVAT.ai, "COCO (CVAT format documentation)". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
  4. Northcutt, Athalye and Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  5. Ultralytics, "SKU-110K dataset documentation". https://docs.ultralytics.com/datasets/detect/sku-110k
  6. arXiv:2006.12634, "Retail product recognition paper" (2020). https://arxiv.org/pdf/2006.12634v1
  7. arXiv:2305.13261, "A Review of Benchmarks for Visual Defect Detection in the Manufacturing Industry" (2023). https://arxiv.org/pdf/2305.13261
  8. National College of Ireland (NORMA eRepository), "Car damage detection thesis". https://norma.ncirl.ie/6101/1/shubhamsarjeraochaudhari.pdf
  9. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  10. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  11. arXiv:2506.17185, "A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset" (2025). https://arxiv.org/pdf/2506.17185
  12. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  13. Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
  14. Carlini et al., "Extracting Training Data from Diffusion Models" (2023). https://www.usenix.org/conference/usenixsecurity23/presentation/carlini
  15. IEEE DataPort (Electric Power Research Institute), "EPRI Distribution Inspection Imagery". https://ieee-dataport.org/open-access/drone-based-distribution-inspection-imagery
  16. Hukkelås and Lindseth, "Does Image Anonymization Impact Computer Vision Training?" (2023). https://openaccess.thecvf.com/content/CVPR2023W/WAD/papers/Hukkelas_Does_Image_Anonymization_Impact_Computer_Vision_Training_CVPRW_2023_paper.pdf
  17. Adobe, "Model release (Adobe Stock contributor help)". https://helpx.adobe.com/stock/contributor/legal/model-release.html
  18. pocstock, "Model release". https://pocstock.com/legal/model-release
  19. WebDataset project, "webdataset (GitHub repository)". https://github.com/webdataset/webdataset
  20. arXiv:2209.15314, "Did You Get What You Paid For? Rethinking Annotation Cost of Deep Learning Based Computer Aided Detection in Chest Radiographs" (2022). https://arxiv.org/pdf/2209.15314
  21. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  22. European Parliament and Council of the European Union, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
  23. NeurIPS, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
  24. NEMA / DICOM Standards Committee, "DICOM PS3.15 Security and System Management Profiles, Annex E: Attribute Confidentiality Profiles". https://dicom.nema.org/medical/dicom/current/output/chtml/part15/chapter_E.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data