Skip to content

Data licensing for AI training

Licensing images for AI training: stock, editorial and commissioned photos

Quick answer

To license images for AI training, you need an explicit grant to use the images as training input, not a standard stock license written for publishing a picture in an ad or article. Then check four layers the copyright license does not cover: model releases for recognizable people, property releases and trademarks in frame, editorial-only restrictions, and rights to captions and metadata. Commissioned shoots can solve all four at once if the contract assigns or licenses training use and collects releases at capture.

By SourceX Editorial · Updated

This guide compares the three main image sources, explains which terms matter, and gives a clause checklist your counsel can mark up. It sits in the data licensing for AI training hub and stays focused on license terms; for the record types themselves see licensing images and inspection photos.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why a standard stock license does not cover training

A standard royalty-free stock license grants the right to reproduce and display an image in a creative work, so training a model is usually outside its scope unless the agreement says otherwise. Typical stock end-user license agreements (EULAs) are drafted around "use in a project": a seat count, print-run caps, and prohibitions on redistribution or use "as a standalone file." Training copies every image into a pipeline, transforms it, and produces weights that can be distributed far beyond any one project.

Some stock EULAs also add explicit machine-learning prohibitions, and large libraries have sold separate data or AI-training licenses to model developers instead [9]. That split tells you how the market reads the base license: training is a distinct right that is priced and scoped separately. Treat any image you bought for marketing as unlicensed for training until the paper says otherwise.

A usable training grant names the activity ("use as input to train, fine-tune, evaluate and test machine learning models"), the model types covered, and what happens to the weights. For drafting language, see writing the AI training rights grant.

Stock, editorial and commissioned images compared

The three sources differ mostly in what releases exist and who can grant training rights, not in image quality. Stock creative content usually comes with model and property releases because contributors must supply them before sale. Editorial content is licensed for newsworthy or descriptive use precisely because releases were never collected. Commissioned work lets you design the rights in from the start.

Illustrative example: invented to show structure; it does not describe an available dataset.

FactorStock creativeEditorialCommissioned shoot
Training right in base licenseUsually no; needs AI/data licenseUsually no; often expressly barredYes, if contract grants or assigns it
Model releasesTypically on file for recognizable peopleTypically absentCollected at capture; can name AI training
Property/trademark clearanceLogos often retouched outLogos, venues and artworks common in frameControlled by shot list
Captions and keywordsContributor or library authored; rights varyJournalist captions with names, places, datesYou own or license them under the contract
Main riskScope of training grant and output termsPublicity rights and identifiable peopleContract gaps: assignment, releases, crew IP
Good fitBroad generative pre-trainingRarely suitable; narrow research with counsel sign-offDomain-specific CV and fine-tuning sets

For commissioned work, decide early between assignment and license; commissioned data: IP assignment vs license covers the tradeoffs, and buying data outright vs licensing it covers the broader question.

Model releases: what they cover and what they miss

A model release is the subject's permission to use their likeness, and it is required whenever a person is recognizable, including by distinctive features such as tattoos rather than the face alone [1]. Most legacy releases were written for advertising and editorial use. Whether "any lawful purpose" wording extends to training a model, or to a generator that can output a near-likeness, is a question for counsel, so ask suppliers for the release template version attached to each asset.

The risk is concrete because image models can memorize. Carlini et al. extracted more than a thousand training images from diffusion models, including photos of identifiable individuals [2]. If a model can regenerate a subject's photo, the subject's release terms matter for outputs, not only for training copies.

Biometric statutes add a separate layer. Face images used to build face recognition or embedding features can trigger laws such as Illinois BIPA, and the FTC's Everalbum order shows the remedy regulators can reach for: deletion of the photos and of the models and algorithms developed from them [3]. That makes release scope and consent records a model-retention issue, covered in what happens to trained models when a data license ends.

Editorial-only images and why they rarely fit

Editorial-only images are the weakest training source because the license, the absence of releases and the subject matter all point the same way. Editorial licenses typically limit use to newsworthy or public-interest contexts and forbid commercial, promotional or merchandising use. A commercial model trained on press photos does not obviously fall within either.

Editorial collections also concentrate the content that creates downstream exposure: celebrities, athletes in branded kit, stadium signage, artworks in galleries and crowds of identifiable bystanders. If your team still wants editorial material, for example for a news-image captioning model, get a written training grant that names the use, a carve-out for publicity claims, and filtering requirements before ingestion.

Trademarks, logos and artworks in frame

Copyright in the photograph is only one of the rights an image can carry; logos, product designs, architecture and artworks in frame are separate rights the photographer usually cannot license. Stock platforms ask contributors for property releases when private property, artwork or branded products are the focus. Diffusion models have been shown to reproduce trademarked logos from training data [2], which turns background clutter into an output risk.

Practical controls for buyers:

  • Ask whether the supplier ran logo detection or retouching, and get the flag rate per batch.
  • Require a property_release_id or "no release required" flag at asset level, not at collection level.
  • Exclude artworks, murals and sculpture in frame unless a property release exists.
  • Keep a deny-list of marks and test generators for logo regurgitation before release.

Warranties and indemnities should match this residual risk; see data warranties for AI training licenses and IP indemnities for licensed training data.

Captions, keywords and embedded metadata

Captions and metadata are often worth as much as the pixels for text-to-image and multimodal training, and they may be licensed separately or not at all. Library keywords and descriptions can be authored by the library, the contributor or a vendor tool, and the image license may not cover them. Confirm in writing that captions, alt text, keywords and category labels are included in the grant.

Embedded metadata is also a privacy channel. An audit of a large web-scraped image set found Exif tags carrying timestamps, geolocation and personal names, and noted that the dataset's download tool extracts these tags for every sample [6]. Specify which Exif, IPTC and XMP fields must be stripped (GPS coordinates, camera serial numbers, creator names) and which you want kept (capture date, lens, orientation).

Open datasets show why per-layer rights tracking matters. COCO's annotations are released under CC BY 4.0, while the underlying images remain subject to the terms their Flickr owners chose [8]. Audits of dataset hosting sites have found license information frequently missing or wrong [7], so verify the image layer and the annotation layer independently; the open data license compatibility matrix helps for public sets.

Regulatory hooks on image training rights

For general-purpose AI models placed on the EU market, text-and-data-mining rights reservations attached to images must be identified and honored. Article 53(1)(c) of the EU AI Act requires general-purpose AI model providers to keep a copyright policy that identifies and complies with text-and-data-mining opt-outs under Article 4(3) of Directive (EU) 2019/790 [4]. As of October 2026 these duties have applied since 2 August 2025, with AI Office enforcement for new models from 2 August 2026. The GPAI Code of Practice copyright chapter describes how signatories can show compliance [5].

For licensed images, this means asking suppliers how opt-outs were checked at collection and keeping that evidence with the license. Image-specific case law on fair use and training remains unsettled as of October 2026, so do not rely on a litigation outcome to fill a gap in the license.

Image license clause checklist

Use this checklist to mark up any image license before ingestion. It covers the terms that most often fail in image deals, in the order counsel usually reviews them.

Illustrative example: invented to show structure; it does not describe an available dataset.

#ClauseWhat to requireRed flag
1Training grantTrain, fine-tune, evaluate and test named model types"Use in projects" or silence on ML
2Scope of assetsAsset ID list or manifest with checksums"Collection as updated from time to time"
3Captions and metadataCaptions, keywords, alt text and labels includedPixels only
4Model releasesRelease ID per asset; template version disclosedCollection-level statement only
5Property and marksRelease or "not required" flag per assetNo logo or artwork screening
6Editorial exclusionEditorial-only assets excluded or separately grantedMixed creative and editorial without flags
7OutputsCommercial use of generated outputs permittedOutput restrictions tied to the source library
8Weights after termTrained models survive expiry or terminationDeletion of models on termination
9TakedownsProcess for removing assets and retraining expectationsOpen-ended retraining duty
10Warranties and indemnityTitle, releases and non-infringement warranties"As is" for the release layer
11Opt-out evidenceRecord of TDM reservation checks at collectionNone for EU-facing models
12Metadata hygieneNamed Exif/IPTC/XMP fields strippedRaw files with GPS intact

A per-asset manifest makes clauses 2 to 6 auditable. A minimal record might carry asset_id, sha256, source_type (stock, editorial, commissioned), model_release_id, property_release_id, editorial_only, caption_license, exif_stripped and license_id. The data license negotiation checklist gives fallback positions for each term.

Sourcing images from operating businesses

Operational photos, such as inspection, field-service and condition-grading images held by companies, are a third route alongside stock libraries and commissioned shoots. They come with their own rights questions: who took the photo, whether customers or employees appear in frame, and whether the company's customer contracts allow the release. Pages such as item condition photos for grading models and the image datasets hub cover those record types, and do AI labs buy images? covers demand.

SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and does not source scraped web content or generic CCTV or photos. Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. If your team needs this kind of data, you can describe the images you need on the buyer page.

Find licensed images for AI training

SourceX looks for US businesses that hold the image data you describe, assesses the data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Nothing is contracted until the supplying company agrees, and a request does not guarantee a match. Start an image data request on the SourceX buyer page.

Sources

  1. Adobe Stock Contributor Help, "Model release overview". https://helpx.adobe.com/ca/stock/contributor/content-policies-guidelines/model-property-releases/model-release-overview.html
  2. USENIX Security 2023 (Carlini et al.), "Extracting Training Data from Diffusion Models" (2023). https://www.usenix.org/conference/usenixsecurity23/presentation/carlini
  3. Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
  4. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  5. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
  6. arXiv, "A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset" (2025). https://arxiv.org/pdf/2506.17185
  7. Longpre et al. (Nature Machine Intelligence 6, 2024), "A large-scale audit of dataset licensing and attribution in AI". https://www.nature.com/articles/s42256-024-00878-8
  8. Hugging Face (HuggingFaceM4), "COCO dataset card (README)". https://huggingface.co/datasets/HuggingFaceM4/COCO/blob/refs%2Fpr%2F5/README.md
  9. Shutterstock, "Shutterstock Expands Partnership with OpenAI, Signs New Six-Year Agreement to Provide High-Quality Training Data" (2023). https://www.prnewswire.co.uk/news-releases/shutterstock-expands-partnership-with-openai-signs-new-six-year-agreement-to-provide-high-quality-training-data-301873361.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data