Skip to content

Video data

Licensing Stock Footage for AI Training: What Standard Stock Licenses Do and Don't Cover

Quick answer

A standard stock footage license usually does not let you train a model. Royalty-free and rights-managed licenses are written for putting clips into productions such as ads, films and apps, and the standard terms of several large libraries (such as Adobe Stock and Getty Images) restrict machine learning use unless a separate agreement grants it. To train on stock video you need a dataset license negotiated with the library. Verify that it covers the clip copyrights, the model and property releases, contributor opt-outs and the metadata you plan to use as captions.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What a standard stock license actually grants

A standard stock license grants the right to reproduce a clip inside an end product, not to feed it to a training pipeline. As of October 2026, the standard agreements of several large stock libraries (including Adobe Stock and Getty Images) restrict using content, and in some cases its captions, keywords and metadata, for machine learning or AI purposes unless an invoice or separate license expressly allows it. Read the definition of machine learning use closely, because it may reach fine-tuning, evaluation or uploading clips into third-party AI tools. Libraries that sell training data run a distinct dataset-licensing process for exactly this reason [1].

Three consequences follow for a video-generation or video-language team:

  • Subscription seats are not dataset rights. A team account that downloads 4K clips for marketing cannot redirect those files into a pre-training corpus, even if the clips are already in your storage bucket.
  • Metadata is licensed content too. Titles, descriptions and keyword lists are often the most attractive caption source, and the same restriction can cover them.
  • Allowed AI uses can be narrow. Where a production license mentions AI at all, it may cover only uses such as editing licensed content in AI tools, not model development.

Treat every library's current terms as the controlling text. They change, and a clause you read last year may not be the one attached to this year's invoice.

How stock dataset deals differ from production licenses

A stock dataset deal is a separate, negotiated agreement that defines a corpus, a training use and a delivery, rather than a per-clip license. Libraries that offer this publish a process that runs from inquiry through scoping and a buyer sample stage to contract and delivery [1]. The commercial terms, pricing and allowed downstream uses are set per deal, so compare offers clause by clause rather than by headline volume. For how enterprise data licenses are structured in general, see enterprise data licensing explained.

Compare the two instruments directly:

TermStandard stock licenseStock dataset agreement
UnitSingle clip or download creditDefined corpus (clip IDs, hours, filters)
UseIncorporation into a productionTraining, fine-tuning, evaluation as defined
AI/ML useCommonly restricted; check termsThe point of the contract; scope varies
MetadataLicensed with the clip; often excluded from AI useShould be listed as a delivered field if used for captions
ReleasesDrafted for the production uses the clip is sold forMust be checked against secondary AI use
Contributor opt-outsNot relevantShould define how opted-out clips are excluded
DeliverySingle file downloadBulk transfer with a manifest

For a deeper treatment of clause structure, see AI data license terms explained and the guide to derivative and successor model rights.

Why releases decide whether a clip is usable

A clip is only as trainable as the weakest release attached to it. Stock model releases were historically drafted for advertising and editorial use. Newer releases add explicit consent to secondary use for AI and machine learning training, and older ones may not include it [2]. Ask the library which release version each clip carries and whether clips with legacy releases are filtered out or re-papered.

Video raises the bar on recognizability. Adobe's contributor guidance treats a person as recognizable by face, face, voice, hair, tattoos, clothing or distinctive surroundings, not only by a frontal face [3]. A back-of-head shot with clear dialogue audio can still need a release, so audio tracks deserve the same review as frames.

Property releases matter for interiors, artworks, vehicles and branded products in frame. For the full layering of footage, people, music, brands and on-screen content, see rights layers in a video clip. Music beds and sound effects are frequently licensed separately from the picture and should be stripped or confirmed.

Contributor opt-outs and editorial-only content

Contributor preferences can remove clips from an otherwise licensable library, so the dataset agreement should say how they are applied. Content-credential tools let creators attach machine-readable "do not train" preferences to their work, a signal rather than a technical block [4]. Ask the library whether it honors contributor opt-outs, at what date the exclusion list is frozen, and whether clips that opt out after delivery must be removed from future refreshes.

Editorial-only clips are a separate exclusion. They often lack model or property releases because news use did not require them, and they frequently show identifiable people, logos and private property. Require a manifest field that marks license class and confirm that editorial-only items are excluded from training subsets. If you use this footage only for evaluation, write that scope explicitly; see evaluation-only data license terms.

Regulatory documentation that follows the footage

Your license file is also your disclosure evidence. Under the EU AI Act, general-purpose AI model providers must maintain a copyright compliance policy that identifies and honors rights reservations expressed under Article 4(3) of the DSM Directive and publish a summary of training content, with duties applying since 2 August 2025 [5]. California's AB 2013 required generative AI developers to post training-data documentation, including whether data was purchased or licensed, by 1 January 2026 [6].

In the US, the Copyright Office's Part 3 report on generative AI training remains a pre-publication version as of October 2026, and courts continue to decide fair use case by case [7]. A clean dataset license reduces dependence on how those questions resolve. Record each clip's license source in machine-readable form; the Croissant-RAI vocabulary offers fields for provenance and usage conditions [8].

Request template for a stock video dataset

A precise request gets a precise quote and fewer surprises at sample review. Send libraries a scoped request like the one below, and require a per-clip manifest at delivery.

Illustrative example: invented to show structure; it does not describe an available dataset.

request:
  use: video-generation pre-training; video-text captioning
  rights_needed: [train, fine_tune, internal_eval]
  content_filters:
    license_class: commercial_only        # exclude editorial-only
    release_version: ai_secondary_use     # exclude legacy releases
    contributor_opt_out: excluded
    audio: strip_music_keep_ambient
  technical:
    min_resolution: 1920x1080
    min_duration_s: 4
    codecs_accepted: [ProRes 422, H.264 high profile]
    deliver_original_frame_rate: true
  metadata_fields_licensed: [title, description, keywords, shot_type, location_country]
  manifest_fields_required:
    - clip_id
    - contributor_id_hashed
    - license_class
    - model_release_id
    - property_release_id
    - release_version
    - opt_out_status_as_of
    - sha256
  sample: 200 clips across top 10 categories before contract
  delivery: WebDataset tar shards, mp4 + json sidecar per clip

WebDataset groups files that share a basename into one sample, so clip_000123.mp4 and clip_000123.json travel together through the loader [9]. For transfer and schema conventions, see dataset delivery formats.

Sample review checklist before you sign

Test the sample against the contract, not against your intuition. Check these items on the sample and again on the first full delivery.

  • Every clip ID in the manifest resolves to a file, and every hash matches.
  • No clip is marked editorial-only, and no clip has a blank release field where people are visible.
  • Release versions match the AI secondary-use requirement for every clip with a recognizable person, including voice-only appearances [2][3].
  • Opt-out status carries a date, and the library states how later opt-outs are handled [4].
  • Keywords and descriptions you plan to use as captions are named as licensed fields.
  • Watermarked previews are not mixed into the delivery, and resolution matches the request.
  • Music beds are stripped or separately licensed.

Duplicates deserve attention too. Stock libraries often hold near-identical takes from the same shoot, which inflates hours without adding diversity and can distort generation training.

When stock footage is the wrong source

Stock video is staged, well lit and centered on marketable subjects, which suits aesthetic pre-training but not every task. Teams building models for operations, procedures or workplace understanding often need real operational footage instead, such as manufacturing assembly video or warehouse operations video. Raw, unpublished takes from production companies are a separate category covered in licensing unpublished creator footage. The video data hub compares sources across the cluster.

SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and does not source scraped web content or generic CCTV. If stock libraries cannot supply the footage your model needs, you can describe the data on the SourceX buyers page.

Sourcing licensed video beyond stock libraries

SourceX looks for US businesses that hold the video you describe, reviews ownership and consents, and delivers each dataset under a license that defines records, uses, term and delivery, only after the supplying company approves the release. Data is sourced on request rather than held in stock, so a request does not guarantee a match, and terms are agreed per deal. Start a video data request at sourcex.si/buyers.

Sources

  1. Pocstock, "The dataset licensing process: from inquiry to delivery". https://support.pocstock.com/en/articles/14772883-the-dataset-licensing-process-from-inquiry-to-delivery
  2. Pocstock, "Model Release". https://pocstock.com/legal/model-release
  3. Adobe, "Model release (Adobe Stock contributor legal guidance)". https://helpx.adobe.com/stock/contributor/legal/model-release.html
  4. MIT Technology Review, "Adobe wants to make it easier for artists to blacklist their work from AI scraping" (2024). https://www.technologyreview.com/2024/10/08/1105234/adobe-wants-to-make-it-easier-for-artists-to-blacklist-their-work-from-ai-scraping
  5. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  6. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  7. U.S. Copyright Office, "Artificial Intelligence Study". https://copyright.gov/policy/artificial-intelligence/
  8. Jain et al. (MLCommons), arXiv, "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  9. WebDataset project (GitHub), "webdataset". https://github.com/webdataset/webdataset

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data