Skip to content

Video data

Rights Layers in a Video Clip: Footage, People, Music, Brands and On-Screen Content

Quick answer

Video rights clearance for AI training means confirming, clip by clip, that the licensor can grant training use for every layer inside the frame and soundtrack: the footage copyright, each recognizable person, every voice, the music composition and the sound recording, visible artwork and trademarks, and on-screen software or documents. Owning the camera file clears only the first layer. Ask for evidence per layer, record it per clip in the manifest, and decide in writing what happens to clips that fail a layer.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why one clip carries seven rights layers

A single clip is a bundle of separately owned interests, and a training license is only as good as the weakest layer the licensor can actually grant. Stock and archive libraries learned this decades ago for advertising use; AI training adds new questions about whether a model counts as a derivative and whether consent can be withdrawn later [9]. The layers below are the working list most clearance reviews converge on.

  1. Footage copyright: who authored and owns the recorded audiovisual work.
  2. People: likeness and publicity rights of every recognizable individual.
  3. Voices and biometrics: voice as likeness, plus faceprints and voiceprints under biometric statutes.
  4. Music: the composition and the sound recording, which are separate works.
  5. Artwork, architecture and trademarks: murals, posters, logos and branded products in frame.
  6. On-screen content: software interfaces, documents, monitors and broadcast feeds captured incidentally.
  7. Location and context: private property, venue terms and restrictions on where filming happened.

Treat the seventh layer as a working hypothesis rather than settled law; venue and property terms vary widely, so ask about them rather than assuming they do or do not bind a licensee. For the broader cluster context, see the video data hub and the AI data overview.

The footage layer is cleared by showing who authored the recording and how ownership reached the licensor. Under US law copyright vests initially in the author, and for a work made for hire the employer, or a commissioning party under a qualifying signed agreement, is considered the author [2]. A transfer of ownership also needs a signed writing (17 U.S.C. 204(a)). That makes the employment status of whoever pressed record the first fact to pin down.

Common failure modes are concrete. A freelance videographer shot the footage without a written assignment, so the production company holds at most an implied license. A fleet company licenses dashcam video but the camera vendor's terms of service claim rights in uploaded clips. A company bought a library in an acquisition and the asset purchase agreement lists "marketing materials" but not raw rushes.

Request the following evidence for this layer:

  • Employment or contractor agreements with assignment or work-for-hire language for each camera operator or source group.
  • Any assignment or acquisition schedule that moves the library to the current licensor.
  • Platform or device terms (cloud video, dashcam, body camera, conferencing) that may reserve rights or limit reuse.
  • A statement of any prior exclusive licenses that would conflict with training use.

The creator and unpublished footage guide covers production-company chains of title in depth, and stock footage licensing explains why standard stock licenses rarely name training.

People and likeness: when a release is needed

A person needs a release, or a lawful alternative basis, whenever they are recognizable, and recognizability is broader than a visible face. Adobe's contributor rules treat a person as recognizable by face, voice, tattoos, clothing or surroundings [3]. Back-of-head shots in a small workplace, a distinctive uniform with a name badge, or a voice on a team call can all identify someone.

For operational video the people are usually employees, customers or bystanders, and each group needs a different answer. Employees may be covered by workplace notices and consent captured for the purpose, which the worker recording and consent guide walks through. Customers and bystanders usually are not covered by an employer's paperwork, so they are cleared by release, by blurring or by excluding the clip.

Ask whether existing releases mention machine learning, AI training or "any media now known or later developed", and whether they allow sublicensing to a third party. A release written for a marketing video often names only the original producer and the original campaign.

Voices, faces and biometric statutes

Voices and faces raise two separate issues: likeness rights over the person's identity and biometric statutes over measurements derived from them. Tennessee's ELVIS Act (HB 2091, Public Chapter 588) extended the state's publicity protections to voice in 2024 [8]. Illinois BIPA covers scans of face geometry and voiceprints and requires written notice and a written release before collection [5].

As of October 2026, a 2024 amendment treats repeated collection of the same identifier from the same person by the same method as one violation, which narrows damages but not the consent duty [6]. Texas defines biometric identifiers to include voiceprints and records of face geometry and requires notice and consent before capture for a commercial purpose [7]. Whether raw video counts as "capture" of an identifier before any template is extracted is contested, so ask counsel how your pipeline's face detection, speaker diarization or re-identification steps interact with these laws.

Practical requests here: the states where recording took place, whether any biometric notice or consent was collected, and whether the licensor ever ran face or voice templating on the footage. In-cab and meeting video carry the heaviest biometric load; see in-cab driver monitoring video and multimodal meeting recordings.

Music: composition and sound recording are two rights

Music in a clip is two copyrighted works, not one: the musical composition (music and lyrics) and the sound recording of a particular performance. US copyright law treats them as distinct works: copyright in a sound recording does not cover the underlying composition, and the authors and owners can differ. A licensor who owns its footage almost never owns either music right for background radio, a retail store playlist or a ringtone on a factory floor.

The usual fix is to strip or replace the music rather than drop the clip. For action recognition, captioning or procedure models, the audio track often matters less than the frames, so a buyer can accept clips with the audio removed or with music segments muted, provided the manifest records that change. Generation teams training audio-visual models need the soundtrack, which makes cleared or commissioned music the only clean path.

Ask suppliers to flag clips with detected music, state whether audio was removed or muted, and supply any sync or master licenses they rely on with the training use spelled out.

Artwork, trademarks and on-screen content

Third-party content visible in the frame is a separate layer, cleared by property releases, by exclusion or by masking. Stock market practice requires property releases for identifiable private property, recognizable artwork and trademarked objects [4]. Logos also matter technically: diffusion models have been shown to memorize and regenerate training images, including trademarked logos and photos of individual people [13].

On-screen content is the layer most often missed in operational video. Screencasts, surgical displays, warehouse terminals and meeting recordings capture third-party software interfaces, customer records, emails and licensed documents. Request the supplier's redaction method for screens and documents, and check it against the video anonymization guide and the software screencast guide.

Training signals attached to media should travel with each clip, because a "do not train" flag can signal a rights reservation that your clearance record must account for. Adobe's content credentials let creators attach a preference that their work not be used for AI training [10], C2PA has published clarifications that its technical specification does not contain a standard assertion for training and data mining in its manifests [11], and the Creator Assertions Working Group defines a comparable training-mining assertion [1]. Check for these assertions at ingest and keep the result per clip.

For teams placing general-purpose models on the EU market, Article 53(1)(c) of the AI Act requires a policy to comply with Union copyright law, including identifying and honoring rights reservations under Article 4(3) of the DSM Directive [12]. As of October 2026, these obligations have applied since 2 August 2025, with AI Office enforcement powers applying from 2 August 2026. A per-clip record of opt-out checks gives that policy something to point at.

Evidence checklist and per-clip rights record

The most useful deliverable from a clearance review is a per-clip rights record that a pipeline can filter on, backed by a checklist of evidence requested for each layer. The table below pairs each layer with the evidence and the fallback when evidence is missing.

Illustrative example: invented to show structure; it does not describe an available dataset.

LayerEvidence to requestFallback if missing
Footage copyrightEmployment/contractor assignments, acquisition schedule, device or platform termsExclude the source group
Recognizable peopleReleases naming AI training and sublicensing; employee notice and consentBlur faces and identifying marks, or exclude
Voices and biometricsRecording states; biometric notice and consent; templating historyRemove or replace audio; exclude BIPA-scope clips
MusicSync and master licenses naming training; music-detection logStrip or mute music segments
Artwork and trademarksProperty releases; logo-detection logMask or crop; exclude for generation use
On-screen contentScreen and document redaction method and QA sampleMask regions; exclude
LocationVenue or property permissions where relevantExclude private-venue clips
Opt-out signalsC2PA or other training assertions checked at ingestExclude flagged clips

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "clip_id": "wh-0412-cam03-000187",
  "source_group": "site-B fixed cameras",
  "footage_owner_basis": "work_for_hire_employee_operated",
  "people": {"recognizable_count": 2, "basis": "employee_consent_v3", "bystanders_blurred": true},
  "audio": {"status": "music_muted", "segments_muted_s": [[12.0, 41.5]], "speech_present": true},
  "biometric_states": ["IL"],
  "biometric_consent": "written_release_on_file",
  "third_party_marks": {"logos_masked": 1, "screens_redacted": 0},
  "tdm_assertion_checked": true,
  "tdm_assertion_value": "none_found",
  "allowed_uses": ["pretraining", "sft", "evaluation"],
  "excluded_uses": ["video_generation"],
  "withdrawal_flag": false
}

Fields such as allowed_uses and withdrawal_flag make the license enforceable in practice: your data loader can drop clips whose permitted uses do not match a training run. The video packaging and sidecar metadata guide shows where this record sits next to clip indexes, and data provenance explains the lineage terms.

Contract terms that close the gaps

Clearance evidence only protects a buyer if the license ties it to the delivered clips and defines what the model is. The archive disputes summarized by Terms.law trace back to three omissions: undefined training scope, unclear derivative status of the trained model, and no mechanism for consent withdrawal [9]. Write all three down.

  • Scope: list the permitted uses (pre-training, fine-tuning, evaluation, generation) and whether outputs may resemble identifiable people, marks or music.
  • Model status: state whether trained weights are or are not derivatives of the licensed clips and what survives termination.
  • Withdrawal: define how a revoked release or a new opt-out flag is notified, by clip_id, and what the buyer must do with the clip and its derived features.
  • Representations per layer: ask the licensor to represent which layers it cleared and to disclose which it did not.

For term-by-term language, see AI data license terms explained and the training data due diligence checklist. Clips combined from several rightsholders follow the pattern in licensing multimodal records from several rightsholders.

How SourceX handles rights in sourced video

SourceX sources operational datasets, including new recordings of hands-on work, from US companies on request, and every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, and every release is approved by the supplying company. You can describe the video data you need and the layers your use case depends on; see also licensing video recordings for AI training.

Request rights-reviewed video data for AI training

SourceX looks for US businesses that hold the video you describe, assesses data and licensing permissions, and agrees allowed uses in a license before anything is transacted. Nothing is contracted until a supplier agrees, and a request does not guarantee a match. Start a video data request.

Sources

  1. IPTC Metawatch, "CAWG Training and Data Mining Assertion". https://metawatch.iptc.org/ai-policy/cawg-training-mining/
  2. U.S. Government Publishing Office (govinfo), "17 U.S.C. 201 - Ownership of copyright (2024 edition)" (2024). https://www.govinfo.gov/content/pkg/USCODE-2024-title17/html/USCODE-2024-title17-chap2-sec201.htm
  3. Adobe Stock Contributor Help, "Model release". https://helpx.adobe.com/stock/contributor/legal/model-release.html
  4. Pocstock, "Model release". https://pocstock.com/legal/model-release
  5. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://ilga.gov/Legislation/ILCS/Articles?ActID=3004&ChapterID=57&Print=True
  6. Illinois General Assembly, "Public Act 103-0769" (2024). https://ilga.gov/legislation/publicacts/103/103-0769.htm
  7. Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  8. Tennessee General Assembly, "HB 2091 (ELVIS Act) bill history" (2024). https://wapp.capitol.tn.gov/apps/Billinfo/default.aspx?BillNumber=HB2091&ga=113
  9. Terms.law, "AI arbitration of archive imagery licensing (case study)". https://terms.law/case-studies/ai-data-licensing-archive-imagery.html
  10. MIT Technology Review, "Adobe wants to make it easier for artists to blacklist their work from AI scraping" (2024). https://www.technologyreview.com/2024/10/08/1105234/adobe-wants-to-make-it-easier-for-artists-to-blacklist-their-work-from-ai-scraping
  11. C2PA (Linux Foundation project), "C2PA clarification to C2PA TDM assertions reference". https://c2pa.org/c2pa-clarification-to-c2pa-tdm-assertions-reference/
  12. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  13. USENIX Security 2023 (Carlini et al.), "Extracting Training Data from Diffusion Models" (2023). https://www.usenix.org/conference/usenixsecurity23/presentation/carlini

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data