Skip to content

Multimodal and embodied data

Licensing Multimodal Records Assembled from Several Sources and Rightsholders

Quick answer

Multimodal dataset licensing works only when rights are mapped per component, not per dataset. A joined record, such as a sales call with its transcript, shared-screen capture and CRM opportunity, can have one owner for the recording system, different people whose faces and voices appear, third-party software and documents on screen, and customer-owned files attached. Buyers should require a component-level rights map, explicit AI-training language in releases and consents, per-asset signals where they exist, and a license that names permitted uses per modality.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why one dataset-level license rarely covers a joined record

A single grant is insufficient because the licensor usually owns the container, not every work and likeness inside it. An audit of widely used public datasets found that datasets assembled from several sources carry obligations that differ by component, so a top-level license can overstate what a commercial user may do [6]. The same failure shows up in enterprise records, where the system of record and the content it stores answer to different parties.

Typical joins that create layered rights:

  • Call plus screen plus CRM. A Gong- or Zoom-style recording, its ASR transcript, the screen share showing a vendor's SaaS UI, and the Salesforce opportunity it links to. See multimodal meeting recordings.
  • Robot log plus facility video. ROS bag trajectories owned by an operator, overhead cameras owned by the facility, and workers in frame. See who owns robot data.
  • CAD plus customer specification. A supplier's STEP or SolidWorks files paired with a customer's drawing package and email thread.
  • Support ticket plus attachments. End-user screenshots and PDFs inside a Zendesk or Jira record. See support tickets with screenshots.

For a single clip's layers (footage, people, music, brands), the video clip rights layers guide goes deeper; this page covers records joined across systems and owners. The multimodal training data hub lists related sourcing guides.

Who holds which rights in a multimodal record

Each component has its own rights holder, and the buyer's job is to name every one before price is discussed. Four layers recur across nearly every joined record.

  1. System and data owner. The company that operates the recording platform, CRM, MES or robot fleet usually controls the export, but its vendor contracts may restrict data reuse.
  2. Recorded people. Employees, customers, contractors and bystanders whose faces, voices or names appear. Recognizability is broader than a face: stock-platform policy treats voice, tattoos, clothing and even distinctive surroundings as grounds for requiring a release [1].
  3. Embedded third-party works. Software UIs, slide decks, licensed stock imagery, background music and brand marks captured on screen or in frame. See screen recordings and third-party content and third-party content inside licensed corpora.
  4. Customer-owned material. Specifications, contracts and files a supplier's own customers sent in. Whether the supplier may license these depends on its customer agreements, and a supplier that quietly expands its terms to permit AI training can draw scrutiny; FTC staff have warned that adopting more permissive data practices this way may be unfair or deceptive [12].

Releases and consents: what language actually covers AI training

A release covers AI training only if it says so, and older releases often do not. Some current model releases name artificial intelligence and machine-learning training as a covered use explicitly [2]; a 2015 marketing-video release that grants "advertising and promotional use" does not obviously extend to training a generative model.

Voice and face data add statutory layers. Illinois BIPA defines biometric identifiers to include voiceprints and scans of face geometry and requires written releases and a published retention schedule, while excluding photographs themselves [7]. Texas requires notice and consent before capturing a biometric identifier for a commercial purpose [8], and California Penal Code 632 requires all-party consent to record a confidential communication [9]. For the Illinois specifics, see the BIPA guide for AI data licensing.

When releases are missing or ambiguous, the practical choices are to drop the modality, redact it, or obtain new consent. Redaction across faces, voices and on-screen text is its own discipline; see de-identifying multimodal records.

Per-asset AI-use signals and how to check them on delivery

Per-asset signals exist for some media and should be read on every delivered file that carries them. The Creator Assertions Working Group defines a cawg.training-mining assertion that sits in a C2PA manifest and marks ai_training, ai_generative_training, ai_inference and data_mining as allowed, notAllowed or constrained [3]. C2PA has clarified that this assertion is referenced from the CAWG specification rather than defined in the core label set [4].

These signals express a preference; they do not technically prevent use, and their force depends on recipients honoring them [5]. For EU-facing general-purpose model providers, Article 53(1)(c) of the AI Act requires a copyright policy that identifies and complies with rights reservations under Article 4(3) of the DSM Directive [10], and the GPAI Code of Practice copyright chapter sets out how signatories implement that policy [11]. As of October 2026, Article 53 obligations apply and AI Office enforcement powers for new models began on 2 August 2026.

A delivery-time check should parse manifests with a C2PA validator (for example c2patool), log the assertion per asset ID, and quarantine files marked notAllowed or constrained until counsel reviews them. Absence of a manifest is not permission; it only means the rights question falls back to the license and releases.

Component rights map: the artifact to request before you sign

A component rights map is a table, one row per modality in the joined record, that states owner, people, embedded works, legal basis and permitted use. Ask suppliers to fill it in during diligence and attach it to the license schedule so that permitted uses and takedown handling are stated per modality.

Illustrative example: invented to show structure; it does not describe an available dataset.

ComponentSource system and ownerRecognizable peopleEmbedded third-party contentBasis for AI usePermitted use (proposed)Takedown unit
Call audio (WAV, 16 kHz)Supplier's meeting platform exportSales reps; customer attendeesHold musicEmployee consent with AI clause; customer notice at call startTraining and evaluation, voices replaced or removedPer call ID
ASR transcript (JSON with word timestamps)Supplier, derived from audioNames spoken in textQuoted pricing from partnersDerived from audio; same basisTraining, names pseudonymizedPer call ID
Screen share (MP4)SupplierFaces in video tilesVendor SaaS UI; customer slide deckSupplier rights to own UI onlyEvaluation only until third-party UI reviewedPer segment timestamp
CRM opportunity (Salesforce export, CSV)SupplierContact names, emailsNoneSupplier-owned business recordTraining, direct identifiers removedPer Opportunity ID
Attached proposal (PDF)Supplier's customerSignatory namesCustomer logos and specsCustomer agreement silent on AIExcluded pending permissionPer file hash

Map every row to a stable join key (call ID, Opportunity ID, file hash) so that a withdrawal by one rightsholder can be executed without dropping the whole record. The record-level takedown guide covers the mechanics, and time synchronization checks explain why segment timestamps must survive redaction.

License terms that should change for multi-rightsholder records

The license schedule should describe records, modalities and uses at component level rather than "the dataset." Clauses that buyers commonly negotiate for joined records include:

  • Per-modality use grants. Training, evaluation, retrieval and synthetic-data generation listed separately for each modality, because a screen share cleared for evaluation may not be cleared for generative training.
  • Representation scope. Which rows of the rights map the supplier stands behind, and which (third-party UI, customer files) it explicitly does not.
  • Withdrawal and takedown. The unit of withdrawal, the join key, and how derived artifacts such as transcripts and embeddings are handled.
  • Downstream disclosure. Whether the buyer may describe the data in an EU training-content summary or a California AB 2013 disclosure.
  • Model-provider routing. Whether licensed content may be sent to third-party model APIs for labeling or retrieval; see licensed content and third-party model API terms.

For a field-by-field view of general terms, read AI data license terms explained, and the dataset licensing glossary entry for vocabulary. The broader training data due diligence checklist covers non-rights diligence.

How SourceX approaches layered-rights requests

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, with diligence materials on source, rights, preparation and allowed use prepared per dataset.

Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Every release is approved by the supplying company, delivery runs through private, access-controlled workflows only after an executed agreement, and nothing is contracted until a supplier agrees. Buyers can describe the multimodal records they need on the SourceX buyer page; a request does not guarantee a match.

Clearing multi-rightsholder multimodal data with SourceX

SourceX looks for US businesses that hold the data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before any transfer. Prices are not published; terms are agreed per deal. Start by describing the joined records and modalities you need at sourcex.si/buyers.

Frequently asked questions

Who owns a recording made on a company's meeting platform?

Usually the company that made it controls the file, subject to its platform vendor's terms, but the people recorded and any content shown on screen carry separate rights. Ownership of the recording does not by itself authorize AI training on recognizable faces, voices or third-party works [1].

Does a general marketing model release cover AI training?

Not reliably. Releases that name AI or machine-learning training explicitly exist in current market practice [2]; older, purpose-limited releases should be treated as not covering it until counsel says otherwise.

Is a missing C2PA manifest the same as permission to train?

No. Manifests and cawg.training-mining assertions are optional, so absence only means rights fall back to the license, releases and applicable law [3][5].

Sources

  1. Adobe Stock (Contributor help), "Model release overview". https://helpx.adobe.com/ca/stock/contributor/content-policies-guidelines/model-property-releases/model-release-overview.html
  2. Pocstock, "Model release". https://pocstock.com/legal/model-release
  3. IPTC Metawatch, "CAWG Training and Data Mining Assertion". https://metawatch.iptc.org/ai-policy/cawg-training-mining/
  4. Coalition for Content Provenance and Authenticity (C2PA), "C2PA clarification to C2PA TDM assertions reference". https://c2pa.org/c2pa-clarification-to-c2pa-tdm-assertions-reference/
  5. MIT Technology Review, "Adobe wants to make it easier for artists to blacklist their work from AI scraping" (2024). https://www.technologyreview.com/2024/10/08/1105234/adobe-wants-to-make-it-easier-for-artists-to-blacklist-their-work-from-ai-scraping
  6. arXiv, "Can I use this publicly available dataset to build commercial AI software? Most likely not" (2021). https://arxiv.org/abs/2111.02374v4
  7. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  8. Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  9. California Legislature, "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632
  10. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  11. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
  12. Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data