Image data
Image Datasets for Computer Vision: How AI Teams Source Real-World Images
Quick answer
Image datasets for computer vision come from four places: public benchmarks such as ImageNet, COCO and Open Images; web-scraped collections; licensed photo archives that companies built while doing their work, such as claims, inspection and quality-control photos; and newly commissioned capture. Benchmarks suit prototyping and comparison, but their terms often do not clearly cover commercial training. Archives add real conditions and outcome labels, while commissioned collection adds control over scene, sensor and releases. Many production teams combine the last two.
By SourceX Editorial · Updated
Public, scraped, licensed or commissioned: four routes to vision data
The route decides three things before any pixel is reviewed: whether labels exist, who can grant training rights, and how closely the images match your deployment cameras.
| Route | What you get | Labels | Rights position | Breaks when | Best fit |
|---|---|---|---|---|---|
| Public benchmarks and open datasets | Curated images, published splits | Included, in the benchmark's own taxonomy | Per-dataset terms, often research-only or mixed per image | Domain, sensor or classes differ; commercial use | Prototyping, comparison with published results |
| Web-scraped collections | Volume and visual variety | Alt text, filenames or none | Copyright in every photo, EU opt-outs, personal data in pixels and headers | You need a documented chain of title | Research; hard to defend for commercial training |
| Licensed operational archives | Photos taken during real work and linked to records | Outcome fields in the record; boxes or masks added later | Licensed by the holder, within its customer and staff terms | Capture was inconsistent or the outcome field is noisy | Domain fine-tuning, real-condition evaluation, rare real failures |
| Commissioned collection | New images captured to your protocol | Labeled to your guideline | Assigned or licensed under your contract, with releases you drafted | You need history, outcomes or long-tail variation | Missing classes, new products or sensors, consented images of people |
Synthetic and rendered images are a fifth input, usually for augmentation; the licensed vs synthetic vs scraped data guide compares them.
What public benchmarks can and cannot do for a commercial model
Public benchmarks are built so results can be compared, not so a company can ship a model trained on them, and three gaps appear when a team tries.
Licenses that do not reach every image. A 2021 study of six commonly used public image datasets found potential license-violation risk in five of them if used to build commercial AI software, partly because one dataset can combine sources under different licenses [1]. SeeFar, a geospatial collection, applies CC BY-NC 4.0 to its WorldStrat imagery and CC BY-SA 3.0 IGO to its Sentinel imagery [2]. COCO-format files carry a top-level licenses section beside images and annotations [3], so check which license applies to each image rather than trusting a dataset-level badge.
Unchecked labels. An audit of ten widely used test sets estimated label errors in at least 6% of the ImageNet validation set [4]. Label scope requires checking the files, not just documentation: for example, Ultralytics documentation for SKU-110K clarifies it provides single-class bounding boxes, with no per-SKU category labels [5], agreeing with a 2020 paper that concluded the dataset cannot be used for product recognition [6].
Conditions that are not yours. A review of manufacturing defect benchmarks notes that DAGM is synthetic, holds at most one defect per image and is statistically biased relative to real production data [7]. Benchmarks rarely match your lens, lighting, background or class balance; see camera, lens and lighting diversity.
When public data falls short, teams turn to the web: a National College of Ireland thesis on car damage detection reports that the public dataset was not enough, so the author scraped Google images [8].
Scraped web images carry copyright, opt-out and privacy debt
Scraping solves volume, but every image arrives with its own copyright, opt-out status and often personal data, and the scraper documents none of it.
- Copyright. The U.S. Copyright Office's Part 3 report on generative AI training, a pre-publication version from May 2025, concludes that copying works into training datasets may be prima facie infringement unless an exception such as fair use applies; the report is advisory and does not bind courts [9].
- EU opt-outs. Providers placing general-purpose AI models on the EU market must keep a copyright policy that identifies and complies with rights reservations under Article 4(3) of the DSM Directive, and must publish a summary of their training content; these obligations applied from 2 August 2025 [10].
- Personal data in headers. An audit of DataComp's CommonPool, a web-scraped image-text pool, found Exif tags carrying timestamps, geolocation and details about individuals, many disclosing full names; the dataset's download tool extracts Exif for every sample [11].
- Faces. Illinois' BIPA covers scans of face geometry (not photographs themselves), requires written notice and a written release before collection, and allows $1,000 per negligent and $5,000 per intentional or reckless violation; since a 2024 amendment, collecting the same biometric from the same person by the same method more than once counts as a single violation [12]. The biometric data guide adds the Texas and Washington rules.
- Remedies that reach the model. The FTC's 2021 final order against photo-app developer Everalbum required deleting the models and algorithms built from users' photos and videos [13]. Image diffusion models can also regenerate individual training images, including photos of people and trademarked logos [14].
SourceX does not source scraped web content or generic photos.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Operational photo archives: real outcomes from systems of record
Operational archives are historical photo archives that companies keep to run their business, stored beside the record of what happened. The link is the asset: the photo shows the condition, and the record shows what someone decided about it.
| Source system | Typical images | Label the linked record supplies |
|---|---|---|
| Auto and property claims platforms | First notice of loss (FNOL) and adjuster photos | Damage findings, repair operations, estimate line items, total-loss decisions |
| Inspection and field-service apps | Asset photos tied to a checklist item or work order | Pass or fail findings, defect codes, technician notes |
| Automated optical inspection (AOI) stations | Fixed-camera images of parts and boards | Reject reasons and dispositions from the quality system |
| Drone and aerial inspection programs | Poles, towers, roofs, bridges | Inspection findings and maintenance tickets |
| Retail shelf audits and catalog studios | Shelf photos and packshots | Planogram compliance results, SKU attributes |
| Construction daily logs and punch lists | Progress and QA photos | Issue type, location, closure status |
Industry research bodies sometimes publish this kind of real-world data: the Electric Power Research Institute (EPRI) released about 30,000 drone images of overhead distribution infrastructure, anonymized and with Exif metadata removed [15].
Archives have predictable weaknesses: mixed devices and angles, free-text labels that need normalizing, heavy imbalance toward normal cases, bursts of near-duplicate shots, and bystanders, plates or addresses in frame. Anonymize selectively: in a CVPR 2023 workshop study on detection datasets, traditional blurring or masking hurt training noticeably, especially for whole bodies, while realistic replacement of faces kept the drop minimal [16] (face and plate blurring evidence). Ask who took each photo, since staff, contractors and customers can hold different rights.
SourceX sources operational datasets from US companies on request and manages the commercial process, including the license and ongoing purchases; buyers describe the data, not the businesses, and the supplying company approves every release. Categories such as images and inspection photos are kinds of data SourceX sources, not inventory under contract, and a request does not guarantee a match. Photos are often worth more with the inspection reports that record their findings. Describe the photos and linked records you need, or read how licensing proprietary data works.
Commissioned collection: control over scene, sensor and release
Commissioned collection buys control: you set devices, lenses, lighting, viewpoints, backgrounds, class quotas and release language before the first frame. It is usually the only route to real photos of products not yet on the market or of defects you can stage safely.
It cannot buy history. Staged scenes lack the long-tail variation and outcomes, such as a repair cost or reject decision, that make archive photos valuable for evaluation. With image data collection services, lead time and per-image cost grow with the sites, people and conditions required.
Three contract points decide whether a commissioned set is usable:
- Releases written for AI training. As market practice, Adobe Stock requires a model release whenever a person is recognizable, including by tattoos, clothing or surroundings [17], and pocstock's policy asks for consent to secondary use including AI and machine-learning training [18]. Older releases may not mention training; see model and property releases for AI training images.
- Ownership. Choose between IP assignment and a license in the collection contract.
- A written protocol. Put capture parameters, label guidelines and acceptance tests in the statement of work for custom collection.
Many teams license an archive for outcomes and variation, then commission capture for missing classes, sensors or consented people; see the custom collection versus licensing comparison.
Match the label to the vision task
Buy the least expensive label that still teaches what the model must learn; the task sets the annotation type and the file format.
| Task | Label to buy | Common container | Check before paying |
|---|---|---|---|
| Classification | One or more class labels per image | CSV or folder per class | Taxonomy matches yours; multi-label allowed |
| Object detection | Boxes with a class | COCO JSON or YOLO text files | Box tightness rules; occlusion and truncation flags |
| Instance segmentation | Polygons or masks | COCO polygons or RLE masks [3] | Boundary tolerance; iscrowd handling |
| Keypoints and pose | Named points per object | COCO person_keypoints files [3] | Visibility flags; skeleton definition |
| Anomaly detection | Normal-only training images; masks on test defects | Image folders plus mask PNGs | Test defects come from real production |
| VLM fine-tuning and captioning | Domain text per image | JSONL or WebDataset tar shards [19] | Text describes the image, not the file name |
A chest-radiograph study treats bounding boxes as noisier alternatives to contour annotations [20], so a detector bought with boxes may need masks later. Operational archives add another label source, the outcome field: domain captions from work records shows how adjuster comments and inspection notes become image text, and pre-labeled versus raw images helps decide whether to buy labels or add them.
What each delivered image record should carry
Each delivered image should carry capture metadata, a stated Exif policy, rights fields and label provenance, so reviewers can audit without opening pixels. For high-risk systems under the EU AI Act, Article 10(2) expects data governance covering the origin of data, annotation, labelling and cleaning, bias examination and data gaps [21]. Regulation (EU) 2026/1744 moved high-risk application dates to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems [22].
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"image_id": "img-000418-03",
"file": "shard-000041.tar/img-000418-03.jpg",
"source_system": "auto_claims_photo_store",
"linked_record": {"record_type": "claim", "record_id": "CLM-000418", "split_group": "CLM-000418"},
"capture": {"device_class": "smartphone", "width_px": 4032, "height_px": 3024,
"format": "jpeg", "capture_month": "2024-06",
"viewpoint": "front_left_45", "setting": "outdoor_daylight"},
"exif_policy": {"gps": "removed", "device_serial": "removed", "owner_name": "removed",
"datetime_original": "truncated_to_month", "orientation": "kept"},
"privacy": {"faces_detected": 0, "plates_detected": 1, "plate_action": "blurred",
"deid_method_ref": "DEID-IMG-03"},
"rights": {"photographer_role": "staff_adjuster", "copyright_holder": "supplier",
"recognizable_people": false, "third_party_marks": ["vehicle_badge"],
"license_ref": "LIC-0021", "permitted_uses": ["model_training", "internal_evaluation"]},
"labels": [
{"type": "bbox_xywh_px", "class": "dent", "value": [812, 1404, 610, 388],
"source": "annotation_vendor", "guideline_version": "2.1"},
{"type": "record_outcome", "field": "repair_operation", "value": "replace_front_bumper",
"source": "estimate_line_item"}
]
}
split_group keeps every photo of one claim in the same split, so near-duplicate angles cannot leak from training into test (duplication and contamination checks). exif_policy keeps orientation, which controls how the image renders, and drops GPS and serials (what to strip and keep in Exif). Ask for a datasheet too; NeurIPS 2026 requires Croissant-RAI responsible-AI metadata for its Evaluations and Datasets Track [23]. The delivery formats guide covers shards and transfer.
Start here: image data guides by domain and concern
Find your domain in the table; cross-cutting guides on volume, resolution, annotation and licensing follow it.
| If you are building | Start with |
|---|---|
| Manufacturing visual inspection | Industrial defect images; rare defect coverage; normal-only anomaly sets; weld inspection images |
| Retail and e-commerce vision | Retail shelf images; shelf-to-catalog pairs; product catalog photos; fashion attributes |
| Claims and property models | Vehicle damage photos; property claim estimates and adjuster reports; roof and hail imagery; inspection photos with findings |
| Aerial, drone and satellite models | Aerial vs satellite vs drone; drone utility inspection; satellite imagery licensing; geospatial and location data |
| Construction and site safety | Construction progress photos; PPE and hazard images |
| Agriculture and food | Crop, weed and field images; food images |
| Medical imaging | Sourcing licensed medical imaging; DICOM de-identification |
| Vision-language models | Commercially usable image-caption sets; multimodal training data |
Medical images need extra care: DICOM states that its confidentiality profiles do not guarantee removal of all identifying information [24], and for health records SourceX requires HIPAA de-identification (Safe Harbor or Expert Determination) before anything is considered for a license.
Cross-cutting guides: how many images a model needs, resolution and compression, boxes, polygons, masks or keypoints, annotation guidelines, unlabeled corpora for pre-training and image licensing terms. Document images belong to the document AI hub, moving images to the video hub, and the AI data buyer's guide maps every cluster.
Image sourcing mistakes that surface after training
These mistakes pass a sample review and surface once a model ships.
- Reading the dataset badge, not per-image licenses. Mixed terms carry into commercial weights.
- Splitting by image instead of by object, site or claim. Near-duplicates inflate test scores.
- Stripping every header field. Orientation, capture month and device class support rendering and stratified evaluation.
- Buying volume of easy normal images. Set per-class minimums for the rare failures that matter.
- Testing only on benchmarks. Hold out images from your own deployment cameras.
Looking for real-world images for computer vision?
Describe the task, classes, capture conditions, linked records, volume and allowed uses on the SourceX buyer page. SourceX looks for US companies that hold matching photo archives, checks the data and the supplier's licensing permissions, agrees a license, and coordinates delivery and payment; nothing is contracted until a supplier agrees. Request real-world image data through SourceX.
Guides in this section
- Construction Progress Photo Datasets for Monitoring AIHow to specify jobsite photo datasets for progress monitoring AI: timestamps, plan location, schedule links, photo ownership and worker privacy checks.
- Drone Power Line Inspection Datasets for Utility AIHow to source drone power line and utility pole inspection imagery for AI: component and defect labels, sensor specs, work-order links, CEII review.
- EXIF Metadata in Image Datasets: What to Strip and KeepA field-level policy for EXIF, XMP and IPTC in image training data: remove GPS, serials and names, keep orientation and camera fields for shift analysis.
- Expert Image Captions from Inspection and Claims NotesHow AI teams turn inspection notes, adjuster comments and work-order findings into de-identified image-text pairs for VLM fine-tuning and evaluation.
- Image-Caption Datasets for Commercial Training: Rights MapWhich public image and image-caption datasets allow commercial training, and how to check per-image CC terms, caption rights, attribution and faces.
- Industrial Defect Image Datasets: Benchmarks vs ProductionWhere MVTec AD, DAGM and VISION stop matching real production lines, and how to specify and license labeled defect images for visual inspection models.
- Model and Property Releases for AI Training ImagesWhether model and property releases cover AI training, what secondary-use language to require, and how to handle logos, trademarks and artwork in images.
- Product Catalog Photo Datasets for AI TrainingHow to source licensed product catalog photos linked to attributes for visual search, classification, VLM fine-tuning and catalog enrichment models.
- Rare Defect Images: Fixing Class Imbalance in Inspection AIHow many rare defect images an inspection model needs, and when to pool real defects across plants, generate synthetic ones, or combine both in training.
- Retail Shelf Image Datasets for SKU and Planogram ModelsHow to source retail shelf image datasets for dense SKU detection, out-of-stock and planogram compliance models: public benchmarks, gaps, specs and rights.
- Roof Damage Datasets: Aerial and Hail Imagery for InsurersHow P&C insurance AI teams source roof condition and hail damage imagery: capture modes, ground truth from inspections and claims, labels, licenses.
- Satellite Imagery Licensing for Commercial AI TrainingCheck whether a commercial satellite EULA or Sentinel, Landsat or WorldStrat terms allow model training, and how derived products and model ownership work.
- Vehicle Damage Datasets for Claims and Repair AI ModelsHow to source licensed vehicle damage photos with part, damage-type and severity labels linked to claim outcomes, beyond small scraped public sets.
- Aerial vs Satellite vs Drone Imagery for Vision ModelsChoose satellite, aerial or drone imagery by target size: map ground sample distance, revisit and viewing angle to what your vision model must resolve.
- Agricultural Image Datasets for Crop and Weed ModelsHow to evaluate and license agricultural image datasets for weed detection and spot spraying: labels, field metadata, coverage gaps and commercial terms.
- Bounding Box vs Polygon vs Mask vs Keypoint AnnotationChoose between bounding boxes, polygons, segmentation masks and keypoints when buying image labels: task fit, cost drivers, COCO fields and spec template.
- Buying Licensed Medical Imaging Datasets for AI TrainingHow to buy medical imaging data for AI: sourcing channels, HIPAA pathways, DICOM de-identification, data use agreements and the metadata to require.
- Camera and Lighting Diversity in Image DatasetsHow to specify camera, lens, lighting and site diversity in an image dataset purchase, record capture metadata, and split evals to measure domain shift.
- Concrete Crack and Bridge Defect Image Datasets for CVHow to source concrete crack, spalling and rust image data for bridge and structure inspection models: labels, capture, imbalance, ownership and licensing.
- Construction PPE Detection Datasets: Hard Hats and HazardsHow to source construction PPE and hazard detection image datasets: class taxonomy, site and camera diversity, label QA, worker privacy and license checks.
- Corrosion Detection Datasets for Asset Integrity AIHow to source corrosion and coating-failure images for segmentation and severity grading: rust scales, label schemas, asset coverage and work-order links.
- DICOM De-identification for AI Training: Annex E and PixelsWhich DICOM PS3.15 Annex E profile and options to require for AI training data, plus private tags, UID remapping, burned-in pixel PHI and residual risk.
- Fashion Image Datasets with Fine-Grained Attribute LabelsHow to source fashion and apparel image datasets with neckline, sleeve and pattern labels, separated views, and rights that allow commercial AI training.
- How Many Images to Train a Computer Vision Model?Size an image dataset per class and per capture condition, use transfer learning and learning-curve pilots, and buy images in stages, not one bulk order.
- Item Condition Grading Photos for Returns and Resale AIHow to specify item condition grading image data: grade taxonomies, defect labels, disposition outcomes, grader agreement and return-photo privacy.
- Normal-Only Image Sets for Anomaly Detection TrainingHow to specify a normal-only training set and labeled anomaly test set for unsupervised visual inspection: coverage, contamination checks, split design.
- Pre-Labeled vs Raw Image Datasets: What to BuyDecide whether to buy pre-labeled images, license raw images and annotate them, or combine both, with a cost model, label audit steps and a decision table.
- Property Inspection Photos Dataset Paired With FindingsHow to source property inspection photos paired with inspector findings for underwriting and claims models: schema, labels, in-home privacy and licensing.
- SKU Recognition Data: Shelf Crops and Catalog ReferencesWhat a detect-then-recognize SKU pipeline needs: reference galleries per SKU, shelf crops linked to catalog images, variant coverage and version history.
- Weld Defect Image Datasets: Visual and Radiographic DataHow to source labeled weld defect images for inspection models: modalities, ISO 6520-1 taxonomy, class imbalance, licensing limits and a request spec.
Sources
- arXiv:2111.02374, "Can I use this publicly available dataset to build commercial AI software?" (2021). https://arxiv.org/abs/2111.02374v4
- Registry of Open Data on AWS, "SeeFar". https://registry.opendata.aws/seefar/
- CVAT.ai, "COCO (CVAT format documentation)". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
- Northcutt, Athalye and Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Ultralytics, "SKU-110K dataset documentation". https://docs.ultralytics.com/datasets/detect/sku-110k
- arXiv:2006.12634, "Retail product recognition paper" (2020). https://arxiv.org/pdf/2006.12634v1
- arXiv:2305.13261, "A Review of Benchmarks for Visual Defect Detection in the Manufacturing Industry" (2023). https://arxiv.org/pdf/2305.13261
- National College of Ireland (NORMA eRepository), "Car damage detection thesis". https://norma.ncirl.ie/6101/1/shubhamsarjeraochaudhari.pdf
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- arXiv:2506.17185, "A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset" (2025). https://arxiv.org/pdf/2506.17185
- Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
- Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
- Carlini et al., "Extracting Training Data from Diffusion Models" (2023). https://www.usenix.org/conference/usenixsecurity23/presentation/carlini
- IEEE DataPort (Electric Power Research Institute), "EPRI Distribution Inspection Imagery". https://ieee-dataport.org/open-access/drone-based-distribution-inspection-imagery
- Hukkelås and Lindseth, "Does Image Anonymization Impact Computer Vision Training?" (2023). https://openaccess.thecvf.com/content/CVPR2023W/WAD/papers/Hukkelas_Does_Image_Anonymization_Impact_Computer_Vision_Training_CVPRW_2023_paper.pdf
- Adobe, "Model release (Adobe Stock contributor help)". https://helpx.adobe.com/stock/contributor/legal/model-release.html
- pocstock, "Model release". https://pocstock.com/legal/model-release
- WebDataset project, "webdataset (GitHub repository)". https://github.com/webdataset/webdataset
- arXiv:2209.15314, "Did You Get What You Paid For? Rethinking Annotation Cost of Deep Learning Based Computer Aided Detection in Chest Radiographs" (2022). https://arxiv.org/pdf/2209.15314
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- European Parliament and Council of the European Union, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
- NeurIPS, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
- NEMA / DICOM Standards Committee, "DICOM PS3.15 Security and System Management Profiles, Annex E: Attribute Confidentiality Profiles". https://dicom.nema.org/medical/dicom/current/output/chtml/part15/chapter_E.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.