Data licensing for AI training
Licensing images for AI training: stock, editorial and commissioned photos
Quick answer
To license images for AI training, you need an explicit grant to use the images as training input, not a standard stock license written for publishing a picture in an ad or article. Then check four layers the copyright license does not cover: model releases for recognizable people, property releases and trademarks in frame, editorial-only restrictions, and rights to captions and metadata. Commissioned shoots can solve all four at once if the contract assigns or licenses training use and collects releases at capture.
By SourceX Editorial · Updated
This guide compares the three main image sources, explains which terms matter, and gives a clause checklist your counsel can mark up. It sits in the data licensing for AI training hub and stays focused on license terms; for the record types themselves see licensing images and inspection photos.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why a standard stock license does not cover training
A standard royalty-free stock license grants the right to reproduce and display an image in a creative work, so training a model is usually outside its scope unless the agreement says otherwise. Typical stock end-user license agreements (EULAs) are drafted around "use in a project": a seat count, print-run caps, and prohibitions on redistribution or use "as a standalone file." Training copies every image into a pipeline, transforms it, and produces weights that can be distributed far beyond any one project.
Some stock EULAs also add explicit machine-learning prohibitions, and large libraries have sold separate data or AI-training licenses to model developers instead [9]. That split tells you how the market reads the base license: training is a distinct right that is priced and scoped separately. Treat any image you bought for marketing as unlicensed for training until the paper says otherwise.
A usable training grant names the activity ("use as input to train, fine-tune, evaluate and test machine learning models"), the model types covered, and what happens to the weights. For drafting language, see writing the AI training rights grant.
Stock, editorial and commissioned images compared
The three sources differ mostly in what releases exist and who can grant training rights, not in image quality. Stock creative content usually comes with model and property releases because contributors must supply them before sale. Editorial content is licensed for newsworthy or descriptive use precisely because releases were never collected. Commissioned work lets you design the rights in from the start.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Factor | Stock creative | Editorial | Commissioned shoot |
|---|---|---|---|
| Training right in base license | Usually no; needs AI/data license | Usually no; often expressly barred | Yes, if contract grants or assigns it |
| Model releases | Typically on file for recognizable people | Typically absent | Collected at capture; can name AI training |
| Property/trademark clearance | Logos often retouched out | Logos, venues and artworks common in frame | Controlled by shot list |
| Captions and keywords | Contributor or library authored; rights vary | Journalist captions with names, places, dates | You own or license them under the contract |
| Main risk | Scope of training grant and output terms | Publicity rights and identifiable people | Contract gaps: assignment, releases, crew IP |
| Good fit | Broad generative pre-training | Rarely suitable; narrow research with counsel sign-off | Domain-specific CV and fine-tuning sets |
For commissioned work, decide early between assignment and license; commissioned data: IP assignment vs license covers the tradeoffs, and buying data outright vs licensing it covers the broader question.
Model releases: what they cover and what they miss
A model release is the subject's permission to use their likeness, and it is required whenever a person is recognizable, including by distinctive features such as tattoos rather than the face alone [1]. Most legacy releases were written for advertising and editorial use. Whether "any lawful purpose" wording extends to training a model, or to a generator that can output a near-likeness, is a question for counsel, so ask suppliers for the release template version attached to each asset.
The risk is concrete because image models can memorize. Carlini et al. extracted more than a thousand training images from diffusion models, including photos of identifiable individuals [2]. If a model can regenerate a subject's photo, the subject's release terms matter for outputs, not only for training copies.
Biometric statutes add a separate layer. Face images used to build face recognition or embedding features can trigger laws such as Illinois BIPA, and the FTC's Everalbum order shows the remedy regulators can reach for: deletion of the photos and of the models and algorithms developed from them [3]. That makes release scope and consent records a model-retention issue, covered in what happens to trained models when a data license ends.
Editorial-only images and why they rarely fit
Editorial-only images are the weakest training source because the license, the absence of releases and the subject matter all point the same way. Editorial licenses typically limit use to newsworthy or public-interest contexts and forbid commercial, promotional or merchandising use. A commercial model trained on press photos does not obviously fall within either.
Editorial collections also concentrate the content that creates downstream exposure: celebrities, athletes in branded kit, stadium signage, artworks in galleries and crowds of identifiable bystanders. If your team still wants editorial material, for example for a news-image captioning model, get a written training grant that names the use, a carve-out for publicity claims, and filtering requirements before ingestion.
Trademarks, logos and artworks in frame
Copyright in the photograph is only one of the rights an image can carry; logos, product designs, architecture and artworks in frame are separate rights the photographer usually cannot license. Stock platforms ask contributors for property releases when private property, artwork or branded products are the focus. Diffusion models have been shown to reproduce trademarked logos from training data [2], which turns background clutter into an output risk.
Practical controls for buyers:
- Ask whether the supplier ran logo detection or retouching, and get the flag rate per batch.
- Require a
property_release_idor "no release required" flag at asset level, not at collection level. - Exclude artworks, murals and sculpture in frame unless a property release exists.
- Keep a deny-list of marks and test generators for logo regurgitation before release.
Warranties and indemnities should match this residual risk; see data warranties for AI training licenses and IP indemnities for licensed training data.
Captions, keywords and embedded metadata
Captions and metadata are often worth as much as the pixels for text-to-image and multimodal training, and they may be licensed separately or not at all. Library keywords and descriptions can be authored by the library, the contributor or a vendor tool, and the image license may not cover them. Confirm in writing that captions, alt text, keywords and category labels are included in the grant.
Embedded metadata is also a privacy channel. An audit of a large web-scraped image set found Exif tags carrying timestamps, geolocation and personal names, and noted that the dataset's download tool extracts these tags for every sample [6]. Specify which Exif, IPTC and XMP fields must be stripped (GPS coordinates, camera serial numbers, creator names) and which you want kept (capture date, lens, orientation).
Open datasets show why per-layer rights tracking matters. COCO's annotations are released under CC BY 4.0, while the underlying images remain subject to the terms their Flickr owners chose [8]. Audits of dataset hosting sites have found license information frequently missing or wrong [7], so verify the image layer and the annotation layer independently; the open data license compatibility matrix helps for public sets.
Regulatory hooks on image training rights
For general-purpose AI models placed on the EU market, text-and-data-mining rights reservations attached to images must be identified and honored. Article 53(1)(c) of the EU AI Act requires general-purpose AI model providers to keep a copyright policy that identifies and complies with text-and-data-mining opt-outs under Article 4(3) of Directive (EU) 2019/790 [4]. As of October 2026 these duties have applied since 2 August 2025, with AI Office enforcement for new models from 2 August 2026. The GPAI Code of Practice copyright chapter describes how signatories can show compliance [5].
For licensed images, this means asking suppliers how opt-outs were checked at collection and keeping that evidence with the license. Image-specific case law on fair use and training remains unsettled as of October 2026, so do not rely on a litigation outcome to fill a gap in the license.
Image license clause checklist
Use this checklist to mark up any image license before ingestion. It covers the terms that most often fail in image deals, in the order counsel usually reviews them.
Illustrative example: invented to show structure; it does not describe an available dataset.
| # | Clause | What to require | Red flag |
|---|---|---|---|
| 1 | Training grant | Train, fine-tune, evaluate and test named model types | "Use in projects" or silence on ML |
| 2 | Scope of assets | Asset ID list or manifest with checksums | "Collection as updated from time to time" |
| 3 | Captions and metadata | Captions, keywords, alt text and labels included | Pixels only |
| 4 | Model releases | Release ID per asset; template version disclosed | Collection-level statement only |
| 5 | Property and marks | Release or "not required" flag per asset | No logo or artwork screening |
| 6 | Editorial exclusion | Editorial-only assets excluded or separately granted | Mixed creative and editorial without flags |
| 7 | Outputs | Commercial use of generated outputs permitted | Output restrictions tied to the source library |
| 8 | Weights after term | Trained models survive expiry or termination | Deletion of models on termination |
| 9 | Takedowns | Process for removing assets and retraining expectations | Open-ended retraining duty |
| 10 | Warranties and indemnity | Title, releases and non-infringement warranties | "As is" for the release layer |
| 11 | Opt-out evidence | Record of TDM reservation checks at collection | None for EU-facing models |
| 12 | Metadata hygiene | Named Exif/IPTC/XMP fields stripped | Raw files with GPS intact |
A per-asset manifest makes clauses 2 to 6 auditable. A minimal record might carry asset_id, sha256, source_type (stock, editorial, commissioned), model_release_id, property_release_id, editorial_only, caption_license, exif_stripped and license_id. The data license negotiation checklist gives fallback positions for each term.
Sourcing images from operating businesses
Operational photos, such as inspection, field-service and condition-grading images held by companies, are a third route alongside stock libraries and commissioned shoots. They come with their own rights questions: who took the photo, whether customers or employees appear in frame, and whether the company's customer contracts allow the release. Pages such as item condition photos for grading models and the image datasets hub cover those record types, and do AI labs buy images? covers demand.
SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and does not source scraped web content or generic CCTV or photos. Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. If your team needs this kind of data, you can describe the images you need on the buyer page.
Find licensed images for AI training
SourceX looks for US businesses that hold the image data you describe, assesses the data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Nothing is contracted until the supplying company agrees, and a request does not guarantee a match. Start an image data request on the SourceX buyer page.
Sources
- Adobe Stock Contributor Help, "Model release overview". https://helpx.adobe.com/ca/stock/contributor/content-policies-guidelines/model-property-releases/model-release-overview.html
- USENIX Security 2023 (Carlini et al.), "Extracting Training Data from Diffusion Models" (2023). https://www.usenix.org/conference/usenixsecurity23/presentation/carlini
- Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- arXiv, "A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset" (2025). https://arxiv.org/pdf/2506.17185
- Longpre et al. (Nature Machine Intelligence 6, 2024), "A large-scale audit of dataset licensing and attribution in AI". https://www.nature.com/articles/s42256-024-00878-8
- Hugging Face (HuggingFaceM4), "COCO dataset card (README)". https://huggingface.co/datasets/HuggingFaceM4/COCO/blob/refs%2Fpr%2F5/README.md
- Shutterstock, "Shutterstock Expands Partnership with OpenAI, Signs New Six-Year Agreement to Provide High-Quality Training Data" (2023). https://www.prnewswire.co.uk/news-releases/shutterstock-expands-partnership-with-openai-signs-new-six-year-agreement-to-provide-high-quality-training-data-301873361.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.