Video data
Licensing Stock Footage for AI Training: What Standard Stock Licenses Do and Don't Cover
Quick answer
A standard stock footage license usually does not let you train a model. Royalty-free and rights-managed licenses are written for putting clips into productions such as ads, films and apps, and the standard terms of several large libraries (such as Adobe Stock and Getty Images) restrict machine learning use unless a separate agreement grants it. To train on stock video you need a dataset license negotiated with the library. Verify that it covers the clip copyrights, the model and property releases, contributor opt-outs and the metadata you plan to use as captions.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What a standard stock license actually grants
A standard stock license grants the right to reproduce a clip inside an end product, not to feed it to a training pipeline. As of October 2026, the standard agreements of several large stock libraries (including Adobe Stock and Getty Images) restrict using content, and in some cases its captions, keywords and metadata, for machine learning or AI purposes unless an invoice or separate license expressly allows it. Read the definition of machine learning use closely, because it may reach fine-tuning, evaluation or uploading clips into third-party AI tools. Libraries that sell training data run a distinct dataset-licensing process for exactly this reason [1].
Three consequences follow for a video-generation or video-language team:
- Subscription seats are not dataset rights. A team account that downloads 4K clips for marketing cannot redirect those files into a pre-training corpus, even if the clips are already in your storage bucket.
- Metadata is licensed content too. Titles, descriptions and keyword lists are often the most attractive caption source, and the same restriction can cover them.
- Allowed AI uses can be narrow. Where a production license mentions AI at all, it may cover only uses such as editing licensed content in AI tools, not model development.
Treat every library's current terms as the controlling text. They change, and a clause you read last year may not be the one attached to this year's invoice.
How stock dataset deals differ from production licenses
A stock dataset deal is a separate, negotiated agreement that defines a corpus, a training use and a delivery, rather than a per-clip license. Libraries that offer this publish a process that runs from inquiry through scoping and a buyer sample stage to contract and delivery [1]. The commercial terms, pricing and allowed downstream uses are set per deal, so compare offers clause by clause rather than by headline volume. For how enterprise data licenses are structured in general, see enterprise data licensing explained.
Compare the two instruments directly:
| Term | Standard stock license | Stock dataset agreement |
|---|---|---|
| Unit | Single clip or download credit | Defined corpus (clip IDs, hours, filters) |
| Use | Incorporation into a production | Training, fine-tuning, evaluation as defined |
| AI/ML use | Commonly restricted; check terms | The point of the contract; scope varies |
| Metadata | Licensed with the clip; often excluded from AI use | Should be listed as a delivered field if used for captions |
| Releases | Drafted for the production uses the clip is sold for | Must be checked against secondary AI use |
| Contributor opt-outs | Not relevant | Should define how opted-out clips are excluded |
| Delivery | Single file download | Bulk transfer with a manifest |
For a deeper treatment of clause structure, see AI data license terms explained and the guide to derivative and successor model rights.
Why releases decide whether a clip is usable
A clip is only as trainable as the weakest release attached to it. Stock model releases were historically drafted for advertising and editorial use. Newer releases add explicit consent to secondary use for AI and machine learning training, and older ones may not include it [2]. Ask the library which release version each clip carries and whether clips with legacy releases are filtered out or re-papered.
Video raises the bar on recognizability. Adobe's contributor guidance treats a person as recognizable by face, face, voice, hair, tattoos, clothing or distinctive surroundings, not only by a frontal face [3]. A back-of-head shot with clear dialogue audio can still need a release, so audio tracks deserve the same review as frames.
Property releases matter for interiors, artworks, vehicles and branded products in frame. For the full layering of footage, people, music, brands and on-screen content, see rights layers in a video clip. Music beds and sound effects are frequently licensed separately from the picture and should be stripped or confirmed.
Contributor opt-outs and editorial-only content
Contributor preferences can remove clips from an otherwise licensable library, so the dataset agreement should say how they are applied. Content-credential tools let creators attach machine-readable "do not train" preferences to their work, a signal rather than a technical block [4]. Ask the library whether it honors contributor opt-outs, at what date the exclusion list is frozen, and whether clips that opt out after delivery must be removed from future refreshes.
Editorial-only clips are a separate exclusion. They often lack model or property releases because news use did not require them, and they frequently show identifiable people, logos and private property. Require a manifest field that marks license class and confirm that editorial-only items are excluded from training subsets. If you use this footage only for evaluation, write that scope explicitly; see evaluation-only data license terms.
Regulatory documentation that follows the footage
Your license file is also your disclosure evidence. Under the EU AI Act, general-purpose AI model providers must maintain a copyright compliance policy that identifies and honors rights reservations expressed under Article 4(3) of the DSM Directive and publish a summary of training content, with duties applying since 2 August 2025 [5]. California's AB 2013 required generative AI developers to post training-data documentation, including whether data was purchased or licensed, by 1 January 2026 [6].
In the US, the Copyright Office's Part 3 report on generative AI training remains a pre-publication version as of October 2026, and courts continue to decide fair use case by case [7]. A clean dataset license reduces dependence on how those questions resolve. Record each clip's license source in machine-readable form; the Croissant-RAI vocabulary offers fields for provenance and usage conditions [8].
Request template for a stock video dataset
A precise request gets a precise quote and fewer surprises at sample review. Send libraries a scoped request like the one below, and require a per-clip manifest at delivery.
Illustrative example: invented to show structure; it does not describe an available dataset.
request:
use: video-generation pre-training; video-text captioning
rights_needed: [train, fine_tune, internal_eval]
content_filters:
license_class: commercial_only # exclude editorial-only
release_version: ai_secondary_use # exclude legacy releases
contributor_opt_out: excluded
audio: strip_music_keep_ambient
technical:
min_resolution: 1920x1080
min_duration_s: 4
codecs_accepted: [ProRes 422, H.264 high profile]
deliver_original_frame_rate: true
metadata_fields_licensed: [title, description, keywords, shot_type, location_country]
manifest_fields_required:
- clip_id
- contributor_id_hashed
- license_class
- model_release_id
- property_release_id
- release_version
- opt_out_status_as_of
- sha256
sample: 200 clips across top 10 categories before contract
delivery: WebDataset tar shards, mp4 + json sidecar per clip
WebDataset groups files that share a basename into one sample, so clip_000123.mp4 and clip_000123.json travel together through the loader [9]. For transfer and schema conventions, see dataset delivery formats.
Sample review checklist before you sign
Test the sample against the contract, not against your intuition. Check these items on the sample and again on the first full delivery.
- Every clip ID in the manifest resolves to a file, and every hash matches.
- No clip is marked editorial-only, and no clip has a blank release field where people are visible.
- Release versions match the AI secondary-use requirement for every clip with a recognizable person, including voice-only appearances [2][3].
- Opt-out status carries a date, and the library states how later opt-outs are handled [4].
- Keywords and descriptions you plan to use as captions are named as licensed fields.
- Watermarked previews are not mixed into the delivery, and resolution matches the request.
- Music beds are stripped or separately licensed.
Duplicates deserve attention too. Stock libraries often hold near-identical takes from the same shoot, which inflates hours without adding diversity and can distort generation training.
When stock footage is the wrong source
Stock video is staged, well lit and centered on marketable subjects, which suits aesthetic pre-training but not every task. Teams building models for operations, procedures or workplace understanding often need real operational footage instead, such as manufacturing assembly video or warehouse operations video. Raw, unpublished takes from production companies are a separate category covered in licensing unpublished creator footage. The video data hub compares sources across the cluster.
SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and does not source scraped web content or generic CCTV. If stock libraries cannot supply the footage your model needs, you can describe the data on the SourceX buyers page.
Sourcing licensed video beyond stock libraries
SourceX looks for US businesses that hold the video you describe, reviews ownership and consents, and delivers each dataset under a license that defines records, uses, term and delivery, only after the supplying company approves the release. Data is sourced on request rather than held in stock, so a request does not guarantee a match, and terms are agreed per deal. Start a video data request at sourcex.si/buyers.
Sources
- Pocstock, "The dataset licensing process: from inquiry to delivery". https://support.pocstock.com/en/articles/14772883-the-dataset-licensing-process-from-inquiry-to-delivery
- Pocstock, "Model Release". https://pocstock.com/legal/model-release
- Adobe, "Model release (Adobe Stock contributor legal guidance)". https://helpx.adobe.com/stock/contributor/legal/model-release.html
- MIT Technology Review, "Adobe wants to make it easier for artists to blacklist their work from AI scraping" (2024). https://www.technologyreview.com/2024/10/08/1105234/adobe-wants-to-make-it-easier-for-artists-to-blacklist-their-work-from-ai-scraping
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- U.S. Copyright Office, "Artificial Intelligence Study". https://copyright.gov/policy/artificial-intelligence/
- Jain et al. (MLCommons), arXiv, "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
- WebDataset project (GitHub), "webdataset". https://github.com/webdataset/webdataset
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.