Video data
Film and TV Production Footage for AI Training: Dailies, Plates, Alternate Takes and Talent Terms
Quick answer
Licensing film and TV production footage for AI training means separating two questions: who owns the pictures, and who controls the people and content inside them. The producing company usually holds copyright in dailies, VFX plates and alternate takes, but performer agreements, union AI provisions, state digital-replica laws, licensed music and third-party inserts can each narrow what a training license may grant. The most useful material is often camera-original footage with lens and motion metadata, not the finished master.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Which production assets are worth licensing for video generation
Camera originals and plates usually carry more training signal than delivered masters, because grading, compression and editorial cuts discard information. For a generative video lab, the asset type determines resolution, dynamic range, take diversity and how much of the frame is clean of graphics. The video data hub covers the wider landscape; the table below is specific to scripted and unscripted production.
| Asset | Typical format | Why it matters for training | Common rights complication |
|---|---|---|---|
| Camera originals (OCF) | ARRIRAW, REDCODE R3D, Sony X-OCN, ProRes RAW, or log ProRes 4444 | High bit depth, full sensor, no grade baked in | Performer likeness on every frame; unused takes may fall outside release scope |
| Dailies | Graded ProRes or DNxHD with burned-in timecode and slate | Many takes per setup; natural variation in blocking | Burn-ins and LUTs need removal or a clean-feed alternative |
| VFX plates | OpenEXR sequences, often ACES2065-1 or ACEScg | Clean backgrounds, camera tracks, greenscreen and clean plates | Vendor work product may be owned or licensed under a separate VFX agreement |
| Alternate takes and outtakes | Same as OCF or dailies | Performance variety for motion and expression | Highest performer-consent sensitivity |
| Delivered masters | IMF packages (ProRes or JPEG 2000), broadcast MXF | Final framing and grade, with captions and audio stems | Music, archival clips and stock inserts licensed only for the program |
| B-roll and second unit | Varies | Establishing shots, vehicles, environments, often no principal cast | Location releases, artwork and signage in frame |
For unedited material that a creator or small production company shot outside a union production, see licensing unpublished and raw footage from creators. For library clips sold under standard terms, see what stock footage licenses do and don't cover.
How performer contracts and union AI terms limit training uses
Owning the footage does not settle whether a performer's likeness and performance can be used to train a model, so treat talent terms as a separate diligence track. Guild agreements such as the SAG-AFTRA TV/Theatrical contract added digital-replica and synthetic-performer provisions in 2023, and those provisions focus on how producers create and use replicas inside productions. They do not read as a general grant to license raw footage to a third-party model developer.
As of October 2026, obtain the text of the guild agreement actually in force for each production year, any side letters, and each principal's individual contract, rather than relying on press summaries. Flag any consent the producer obtained for replica creation: it is usually scoped to a project, not to external model training. Alternate takes and outtakes deserve the strictest review, because performers never approved them for release at all.
State law adds another layer. California's AB 2602 makes a contract provision allowing a digital replica of a performer's voice or likeness unenforceable when it replaces work the performer would otherwise have done in person, lacks a reasonably specific description of the intended uses, and the performer was not represented by counsel or a union [1]. Tennessee's ELVIS Act extends that state's right of publicity to voice and likeness, including AI imitations [2], and other states have followed with entertainment-specific rules [3]. A broad, decades-old "all media now known or hereafter devised" clause may not carry the weight a licensor expects. For the voice side of the same problem, see voice talent consent and release terms.
Rights layers inside a single production frame
A production clip often contains several separately owned works, and a footage license only conveys what the licensor holds. Map each layer before you price the deal. The full breakdown sits in rights layers in a video clip; for production footage the recurring layers are:
- Underlying literary rights: options and adaptation rights can restrict derivative uses of characters and story elements.
- Music: sync licenses for score and needle-drops are usually limited to the program; strip or exclude audio stems unless music rights are separately cleared.
- Archival and stock inserts: clips licensed for one program rarely extend to AI training.
- Artwork, set dressing and brands: clearance reports (E&O clearance logs) show what was cleared and for which media.
- Locations and extras: location agreements and background-actor releases vary widely in scope.
- VFX vendor deliverables: CG elements, matte paintings and comps may be vendor-owned or licensed back under the VFX agreement.
The U.S. Copyright Office's Part 3 report, still a pre-publication version as of October 2026, analyzes how training implicates reproduction rights and why licensing markets matter to fair use [4]. Licensed footage with a documented chain of title reduces reliance on contested fair-use arguments.
Camera metadata as conditioning signal
Production metadata can turn footage into structured conditioning data for camera motion, lens and exposure control, so ask for it explicitly. OpenEXR plates can carry camera and lens attributes in their headers, and conform tools can extract per-frame lens data from ARRIRAW, R3D and Sony RAW. Much of the most useful data, though, lives outside the image files: in ALE and EDL files, camera reports, DIT logs and VFX data-wrangler sheets.
Color is the common failure mode. Plates are often exchanged in ACES2065-1 while VFX vendors work in ACEScg, and a dataset that mixes encodings without labeling them will teach a model inconsistent color. Request the color pipeline per shot, not per show.
Fields worth requesting, where they exist:
- Timecode, reel or clip name, scene, take and camera ID from slate or ALE files
- Lens model, focal length, T-stop, focus distance (per frame where lens data systems such as ZEISS eXtended Data or Cooke /i were used)
- Camera body, sensor mode, frame rate, shutter angle, ISO and white balance
- Camera tracking solves (for example, exported from 3DEqualizer or SynthEyes) and LiDAR or survey data for plates
- Color space, transfer function, CDL values and AMF or LUT references
- Shot status: circled take, NG, pickup, or VFX pull
These fields map cleanly to the caption and conditioning schemas discussed in training data for text-to-video models and video-text pairs and dense captions.
Request template for production footage
A precise request describes the footage, the metadata and the rights you need, not a named studio or title. Use this structure when you brief a supplier or broker.
Illustrative example: invented to show structure; it does not describe an available dataset.
request: professional_production_footage
intended_use: video generation pre-training and fine-tuning; camera-motion conditioning
asset_types: [camera_originals, vfx_plates, alternate_takes]
exclude: [delivered_masters_with_music, archival_inserts, burned_in_graphics]
technical:
min_resolution: 3840x2160
formats_accepted: [ARRIRAW, R3D, ProRes 4444 XQ log, OpenEXR]
color: per-shot color space + transfer function declared
frame_rate: native, no pulldown
metadata_required: [timecode, scene, take, camera_id, lens_model, focal_length_per_frame, shutter_angle]
metadata_optional: [camera_track_solve, CDL, AMF]
people_in_frame:
principal_performers: allowed only with documented AI-training consent
background: allowed if releases cover the use, else exclude or blur
rights_evidence: [chain_of_title, performer_consent_records, union_agreement_version, clearance_log]
audio: excluded unless separately cleared
territory_and_markets: model distributed globally, including EU and California users
Cross-border and disclosure obligations that follow the footage
Where your model is placed on the market affects what records you need from the licensor. Under EU AI Act Article 53, general-purpose AI model providers must maintain a policy to comply with Union copyright law, including identifying and respecting rights reservations under Article 4(3) of Directive (EU) 2019/790 [5]. California's AB 2013 requires developers of generative AI systems offered to Californians to post documentation about training datasets, including whether they contain copyrighted material and whether they were purchased or licensed [6].
Japan is often cited as permissive, but commentary on the Agency for Cultural Affairs guidance indicates that Article 30-4 does not cover training aimed at outputs that reproduce the creative expression of specific works [7]. A lab fine-tuning to mimic a particular film's look should read the Agency's own "General Understanding on AI and Copyright in Japan" rather than secondary summaries [8]. For general license structure, see the AI training data licensing guide.
Diligence checklist before signing
A production-footage deal is ready to sign only when each rights layer has a document behind it. Use this as a working list:
- Chain of title from the production company to the licensor, including any completion-bond or financier liens.
- Union agreement version for each production year, plus side letters.
- Principal performer contracts, with AI and digital-replica clauses extracted.
- Background actor and location release templates, with a sample of signed copies.
- Music cue sheets and sync licenses, to decide whether audio is excluded.
- Clearance log for artwork, brands and archival inserts.
- VFX vendor agreements for any plates containing vendor elements.
- Technical manifest: file counts, formats, color spaces, metadata coverage and checksums (for example, MD5 or xxHash from the DIT's offload reports).
Studios approaching this from the supply side face their own compliance checks; the SourceX insight on compliance checks for film and post-production studios covers them, and media and publishing buyers shows how the category fits SourceX's buyer work.
Where SourceX fits
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements. It does not hold stock, so a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, and every release is approved by the supplying company. SourceX does not source scraped web content or generic CCTV or photos. You can describe the footage and metadata you need to SourceX.
Request production footage for AI training
SourceX looks for US businesses that hold the data you describe and handles assessment of data and licensing permissions before any agreement. Nothing is contracted until a supplier agrees, and pricing and allowed uses are set in the license for each deal. Start a buyer request at SourceX.
Frequently asked questions
Do union AI provisions stop a producer from licensing footage for model training?
Guild AI provisions mainly regulate how producers create and use digital replicas and synthetic performers inside productions. Whether a third-party training license is permitted depends on the agreement in force, individual contracts and state law, so counsel should review the actual documents.
Is footage without recognizable performers easier to license?
Usually. Plates, B-roll and second-unit footage without principal cast avoid most performer issues, though location releases, artwork and vendor ownership still apply.
Should we accept dailies with burned-in timecode?
Only if clean versions are unavailable. Burn-ins, watermarks and slate overlays teach models to generate artifacts, so ask for clean camera originals with timecode carried in metadata instead.
Sources
- Cohen Davis Ascher Schochet (CDAS), "California Passes AI Digital Replica Law for Performers" (2024). https://cdas.com/california-passes-ai-digital-replica-law-for-performers/
- AI Law Tracker, "HB 2091 — Tennessee ELVIS Act (AI voice & likeness)". https://ai-law-tracker.com/laws/bill/tennessee-elvis-act
- Davis Wright Tremaine, "Lights, Camera, Legislation: Are Your Entertainment Contracts AI Ready?" (2025). https://dwt.com/insights/2025/03/state-laws-regulating-ai-in-entertainment-industry
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- Privacy World, "Japan's New Draft Guidelines on AI and Copyright: Is It Really OK to Train AI Using Pirated Materials?" (2024). https://www.privacyworld.blog/2024/03/japans-new-draft-guidelines-on-ai-and-copyright-is-it-really-ok-to-train-ai-using-pirated-materials/
- Agency for Cultural Affairs, Government of Japan, "Copyright (English policy page, including General Understanding on AI and Copyright in Japan)". https://www.bunka.go.jp/english/policy/copyright/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.