Image data
Image and Caption Datasets You Can Train On Commercially: Public Benchmarks, CC Collections and Captioned Sets
Quick answer
Few famous image datasets are cleanly usable for commercial training. ImageNet's access terms restrict use to non-commercial research and education [11]; COCO licenses its annotations under CC BY 4.0 but leaves each photo under its Flickr owner's terms [12]. Creative Commons-only collections such as CommonCanvas [2] and newer permissively licensed caption sets [3] are better starting points, but you still need to verify each image's CC variant, the rights to the captions, attribution duties and faces before a commercial run.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why the headline license rarely answers the commercial question
The license on a dataset card usually covers the annotations or the compilation, not the underlying pixels. COCO is the textbook case: the consortium states it does not own image copyright and that image use must follow Flickr's terms, while the boxes, masks and captions are CC BY 4.0. A buyer who reads only "CC BY 4.0" on the card can conclude the wrong thing about the photos.
Mixed-source datasets compound this. Collections stitched together from several upstream corpora inherit each source's terms, and the top-level license may not cover every image [4]. At the ecosystem level, the Data Provenance Initiative found license omission above 70% and license error rates above 50% on popular dataset hosting sites [7]. That audit covered text datasets, but the same hosting platforms serve image-text sets, so treat a Hugging Face "license:" tag as a pointer, not evidence.
For the general, format-agnostic process, see our guide to auditing open dataset licenses before commercial training. This page covers what is specific to images and captions.
Where common image and caption datasets land
Most well-known image benchmarks fall into one of four rights patterns, and each needs a different check. The table below summarizes the patterns as of October 2026; always re-read the current terms of the exact release you download.
| Pattern | Examples | Images | Captions or labels | Commercial training read |
|---|---|---|---|---|
| Research-only access agreement | ImageNet / ILSVRC | Governed by access agreement: non-commercial research and education only; binds a for-profit employer; revocable | Same agreement | Not usable for commercial training without separate rights |
| Split license, per-image source terms | COCO (Flickr-sourced) | Each photo under its own Flickr license and Flickr terms | CC BY 4.0 | Annotations usable with attribution; images need per-image review |
| URL-plus-alt-text pools | LAION-style and CommonPool-style web crawls | Not distributed; you fetch from third-party hosts with unknown rights | Alt text scraped from web pages | High risk: copyright, personal data and stale links [5] |
| CC-filtered or permissively released collections | CommonCanvas [2], GPIC (reported) [3] | Filtered to CC licenses or released under a permissive license | Often synthetic captions generated by a captioning model | Strongest public option; still verify CC variants and caption model terms |
Evaluation sets carry their own terms, which we cover separately in checking public benchmark licenses for commercial use. Note also that benchmark quality is not a given: at least 6% of the ImageNet validation set has label errors [10].
How to read Creative Commons variants at the image level
A "Creative Commons" image is only commercially trainable if its specific variant allows it, and each variant must be checked per file. CC0 and public-domain marks impose no conditions. CC BY and CC BY-SA allow commercial use with attribution, and BY-SA adds share-alike conditions that counsel should map against model weights and any redistributed dataset copy.
CC BY-NC, BY-NC-SA and BY-NC-ND exclude commercial use, and CC-only corpora still have to separate these out. CommonCanvas handled this by building its training set only from CC images and tracking commercial and non-commercial subsets distinctly [2]. Whether training itself is a use that triggers the license conditions is a contested legal question; do not assume either answer without counsel.
Common failure modes in CC image pulls:
- License drift. Flickr owners can change the license shown on a photo page after you scrape it, and a CC grant already made is hard to prove without a record; snapshot the license string and the retrieval date per image.
- License laundering. A photo re-uploaded by someone other than the author carries a CC tag the uploader had no right to grant.
- Version ambiguity. CC 2.0 and CC 4.0 differ on attribution mechanics and sui generis database rights; record the version, not just "CC BY".
- ND confusion. No-derivatives terms matter if you crop, re-encode or composite images into training shards.
Caption rights are a separate question from image rights
Captions have their own provenance, and clearing the image does not clear the text. Alt text and surrounding page text in web-crawled pools were written by site authors who granted nothing [5]. Human captions written for a benchmark take the benchmark's annotation license, as with COCO's CC BY 4.0 captions.
Synthetic captions shift the question to the captioning model. If you re-caption CC images with a vision-language model, as CommonCanvas did [2], check that model's license and acceptable-use terms for restrictions on using its outputs to train other models. Keep the captioner name, version and prompt in your lineage record.
Domain captions are a third route. Inspection notes, adjuster comments and technician write-ups that already sit next to photos in business systems can serve as image text; we cover that pattern in domain captions from work records. Those come with the data owner's rights and consents rather than a public license.
Faces, biometrics and personal data inside image-text pools
Any image corpus with people in it is a personal-data and biometric question, regardless of its copyright license. An audit of a large web-scraped image-text pool found personal data in images, captions and EXIF metadata [5]. A CC license from the photographer does not carry consent from the people photographed. Even ImageNet contains enough incidental faces that researchers studied blurring them across the dataset [1].
MegaFace shows how this fails at scale: it was assembled from Creative Commons Flickr photos and used for face recognition work well beyond what attribution and NC terms allowed [6]. US biometric statutes add a separate layer. Texas, for example, treats a record of face geometry as a biometric identifier and requires notice and consent before capture for a commercial purpose [9]. As of October 2026, the version of that section effective 1 January 2026 adds an exception tied to artificial intelligence systems, reportedly not available where the system is used to identify individuals, so have counsel confirm whether it reaches your training use. Illinois BIPA applies its own consent and retention rules.
Practical controls: run face detection over the pool and route face-bearing images to a separate review bucket, strip GPS and camera serial fields (see EXIF metadata in image training data), and scrub names, emails and phone numbers from captions with an NER pass before tokenization.
Building a per-image rights register
A per-image rights register is the artifact that lets you answer a licensor, an auditor or an EU copyright-policy question later. GPAI providers placing models on the EU market must put in place a copyright policy under AI Act Article 53(1)(c), and the GPAI Code of Practice copyright chapter describes how signatories can show it [8]. As of October 2026, Article 53 duties have applied since 2 August 2025, and AI Office enforcement for new models runs from 2 August 2026.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"image_id": "img_000184233",
"sha256": "9f2c...e41a",
"source_dataset": "cc_filtered_pool_v2",
"source_url": "https://example.org/photo/184233",
"retrieved_at": "2026-09-14",
"image_license": "CC-BY-4.0",
"license_evidence": "license string captured from source page at retrieval",
"attribution": {"author": "J. Example", "required": true},
"nc_flag": false,
"nd_flag": false,
"sa_flag": false,
"caption_source": "synthetic",
"caption_model": "captioner-x v1.3",
"caption_model_terms_checked": true,
"faces_detected": 0,
"exif_stripped": true,
"pii_scrub_caption": "ner_v4, 0 entities",
"eligible_commercial": true,
"exclusion_reason": null
}
Roll the register up to a per-dataset view with license mix, NC share, attribution obligations and face counts. That rollup is what procurement and counsel will sign against, and it also feeds benchmark decontamination; see decontaminating a licensed training set against public benchmarks.
Commercial-use decision checklist for an image-caption set
Run every candidate set through the same gate before it reaches a training shard. The checklist below assumes you have already done the generic license audit.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Check | Pass condition | If it fails |
|---|---|---|
| Access terms | No research-only or non-commercial access agreement | Exclude, or obtain separate commercial rights |
| Image license granularity | License recorded per image with version and retrieval date | Re-crawl license strings or drop unverifiable images |
| NC / ND variants | Filtered out or held in a separate non-commercial pool [2] | Do not merge pools |
| Caption provenance | Human captions under a known license, or synthetic captions from a model whose terms allow it | Re-caption with a permitted model |
| Attribution | Attribution file you can actually publish or ship with the model card | Drop images whose attribution you cannot honor |
| Faces and biometrics | Face-bearing images reviewed; consent basis documented [9] | Blur, drop or replace with consented data |
| Personal data in text and metadata | EXIF stripped, caption PII scrubbed [5] | Hold back until scrubbed |
| Mixed-source inheritance | Every upstream source mapped to its terms [4] | Exclude unmapped sources |
When licensed or commissioned image-text data makes more sense
Public CC pools cover generic scenes well, but they thin out on enterprise domains: equipment interiors, damaged property, defects, medical devices in use, back-office documents photographed in the field. For those, buyers typically license image sets directly from the businesses that hold them, with the text that already accompanies each photo. Our image data hub maps the main categories, and license images and inspection photos for AI training describes that buying route.
SourceX sources operational datasets, including documents and new recordings of hands-on work, from US companies on request and manages the licensing process; it does not source scraped web content or generic CCTV or photos. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. If your captioned-image requirement is domain-specific, describe it to SourceX and specify fields, volume and allowed uses rather than naming companies.
For context, see also do AI labs buy images, our training data due diligence checklist, the glossary entries for multimodal data and data provenance, and video-text pairs and dense captions if you are extending beyond still images.
Source licensed image-caption data for commercial training
SourceX sources operational data, including documents and new recordings of hands-on work, from US companies on request, and each release is approved by the supplying company before delivery under a license. A request does not guarantee a match, and nothing is contracted until a supplier agrees. Start by describing the images, captions and allowed uses you need at sourcex.si/buyers.
Frequently asked questions
Can I use ImageNet to train a commercial model?
Not under its standard terms. The ImageNet access agreement limits use to non-commercial research and educational purposes and extends to a for-profit employer of the researcher. Commercial use requires rights outside that agreement.
Are COCO captions usable commercially?
The COCO annotations, including captions, are licensed CC BY 4.0, so they can be used with attribution. The photos are a separate matter and follow each image's Flickr terms.
Is a CC-only image dataset automatically safe for commercial training?
No. CC-only means you still have to remove NC and ND images, honor attribution, check caption provenance and screen for faces and personal data [2][6].
Sources
- Yang, Russakovsky et al., arXiv / ICML 2022, "A Study of Face Obfuscation in ImageNet" (2021). https://arxiv.org/abs/2103.06191v2
- DAIR.AI Academy (paper summary), "CommonCanvas: An Open Diffusion Model Trained with Creative-Commons Images" (2023). https://academy.dair.ai/papers/commoncanvas
- AI Weekly, "Stanford's GPIC ships 100M commercially-usable training images". https://aiweekly.co/alerts/stanfords-gpic-ships-100m-commercially-usable-training-images
- arXiv, "arXiv:2111.02374v4" (2021). https://arxiv.org/abs/2111.02374v4
- arXiv, "Privacy audit of a web-scraped image-text pool (arXiv:2506.17185)" (2025). https://arxiv.org/pdf/2506.17185
- Exposing.ai, "MegaFace". https://exposing.ai/megaface/
- Longpre et al., arXiv, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier" (2026). https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
- Northcutt, Athalye, Mueller, arXiv / NeurIPS 2021, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Princeton University and Stanford University (image-net.org), "ImageNet Terms of Access" (2026). https://image-net.org/accessagreement
- COCO Consortium (cocodataset.org), "COCO Terms of Use" (2026). https://cocodataset.org/#termsofuse
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.