Skip to content

Data sourcing by buyer team

Reviewing an AI data license as in-house counsel: diligence, risk triage and sign-off

Quick answer

In-house counsel should review an AI training data license in four passes: collect diligence before redlining (source systems, chain of title, the supplier's own customer contracts, consent basis, de-identification method and sample results), triage risk by category with a named escalation owner, negotiate the few clauses that control training use (grant, field of use, model retention, derivatives, warranties, audit and deletion), and close with a sign-off memo that governance and data engineering can execute.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why a data license is not a software license

A training data license fails differently from a software license because the risk travels into the model. Standard SaaS and software terms assume you can stop using the product on termination; with training data, the dataset's influence persists in weights, checkpoints, distilled students and evaluation baselines [9]. Counsel therefore has to review the deal against the model lifecycle, not only against the dataset.

The remedy exposure makes this concrete. The FTC has ordered companies to delete not only improperly obtained data but the models and algorithms trained on it, including in a 2021 photo-app order [1]; a 2023 facial recognition case also required deletion of collected images [3], and FTC technology staff have warned AI companies that using data contrary to their commitments can bring the same result [2]. A provenance gap in an upstream supplier can become a model problem for you.

Copyright posture points the same way. US litigation in 2025 and 2026 has distinguished training sources that were lawfully acquired from copies obtained from pirate sources; in Bartz v. Anthropic, the class settlement received final approval in July 2026 [4]. A written license with a documented chain of title is the cleanest evidence that your acquisition was lawful. For the rights framework behind this, see the buyer's guide to AI training data licensing.

Diligence to request before you redline

Request diligence first, because the right redlines depend on facts the draft license will not tell you. A supplier's first draft usually reflects its own risk appetite; until you know how the records were created, who owns them and what notices covered them, you cannot tell whether a warranty is meaningful or decorative. The training data due diligence checklist lists items to check; the request below is the counsel-specific subset.

Illustrative example: invented to show structure; it does not describe an available dataset.

RequestWhat good looks likeRed flag
Source systemsNamed systems of record (for example Zendesk tickets, Salesforce opportunities, Jira issues, NetSuite journals) with export dates and record counts"Aggregated from multiple sources" with no system list
Chain of titleStatement of how the supplier came to hold each record class: created in its own operations, received from customers, or licensed inRecords originally licensed in from a third party, with no sublicense right [10]
Supplier's own customer contractsRelevant clauses of the supplier's MSA, DPA and terms of service governing use of customer contentDPA limits processing to "providing the services"; no secondary-use or AI clause
Consent and notice basisPrivacy notices in force when records were collected; call-recording consent methodCall audio from California with one-party consent only [12]
De-identificationWritten method (field removal, token replacement, NER-based redaction), tools used, sample size and residual-hit results"Anonymized" with no method or sample test
Third-party content inside recordsInventory of embedded attachments, quoted emails, code from open-source repos, licensed images or articlesPDFs and screenshots attached to tickets, never scanned
Regulated data flagsStatement on PHI, nonpublic personal financial information, biometrics, children's data and export-controlled technical dataSilence on these categories

Two items deserve emphasis. The supplier's contracts with its own customers often decide whether the supplier can license at all; a DPA that makes it a processor for customer data generally means the data is not the supplier's to sell. And embedded third-party content, such as attached contracts, stock images or copied articles inside support tickets, is where otherwise clean operational data most often carries someone else's copyright.

A risk triage grid with escalation owners

Triage each dataset by risk category so that escalation goes to the person who can actually resolve it. Most deals do not need every specialist; the grid below gives a default owner and the trigger that should pull them in.

Illustrative example: invented to show structure; it does not describe an available dataset.

Risk categoryTriggerEscalate toTypical resolution
Rights chainAny record class licensed in or user-generatedIP or commercial counselSublicense evidence, carve-out of the record class, or title warranty with indemnity
Personal dataNames, emails, phones, account numbers, free text about individualsPrivacy counsel or DPODocumented de-identification, sample test, re-identification prohibition
EU or UK data subjectsRecords about people in the EU or UKPrivacy counselAnonymization analysis under GDPR Recital 26 [8], or a lawful basis assessment if personal data remains
Biometrics and recordingsVoice, face, video of peoplePrivacy counsel plus litigationBIPA written release for Illinois residents [11]; all-party consent evidence for California calls [12]
Health dataClinical notes, claims, PHI from covered entitiesPrivacy counsel with HIPAA experienceSafe Harbor or Expert Determination de-identification and the determination documentation
Financial dataCustomer account and transaction data held by a financial institutionRegulatory counselGLBA analysis of whether the supplier may disclose, and de-identification
Confidential third-party informationCustomer names, pricing, contracts of the supplier's clientsCommercial counselRedaction of counterparty identifiers or exclusion
Export-controlled technical dataEngineering records on defense, aerospace or dual-use itemsTrade complianceEAR or ITAR classification before any transfer to foreign-national staff or offshore compute

Severity should track the remedy, not only the probability. A low-probability rights defect that could support a model deletion order deserves more attention than a moderate-probability defect fixable by deleting a few thousand records.

The clauses that matter most for training use

Six clauses decide whether a data license actually supports training; the rest is standard commercial drafting. Spend negotiation capital here, and use the AI data license terms explainer for the vocabulary.

  1. Rights grant and field of use. The grant should name the activities you will perform: pre-training, fine-tuning, evaluation, retrieval-augmented generation, synthetic data generation. A grant limited to "internal analytics" or "research" will not cover a commercial model.
  2. Model retention after termination. Decide whether trained weights, checkpoints and distilled models survive termination or expiry of the license. Silence invites a later argument that the model is a derivative that must be destroyed.
  3. Derivative and successor models. Define whether rights extend to fine-tunes of the trained model, future model versions, and models trained on synthetic data generated from the licensed data.
  4. Warranties of title and authority. Ask for a warranty of title covering ownership or sufficient rights, authority to license for the stated uses, compliance of collection with applicable law and notices, and accuracy of the de-identification description.
  5. Indemnity and its cap. A title warranty without an indemnity that survives and sits outside the general liability cap is weak protection against third-party claims; negotiate a separate cap for IP and privacy claims.
  6. Audit and deletion mechanics. Define what happens when a record is later found defective: notice, record-level deletion from the dataset, whether retraining is required, and how the supplier evidences its own upstream audit.

Practitioner guidance treats provenance, chain of title and sublicensing rights as the central failure points in AI licensing, which is why buyers ask for audit rights over provenance [9][10]. Whether either is market for your deal depends on leverage and the dataset's risk profile.

Disclosure duties that start at signing

Signing a training data license can trigger disclosure duties for your company, so counsel should capture the facts needed for those disclosures before the deal closes. California's AB 2013 requires developers of generative AI systems offered to Californians to post documentation about training data, including sources and whether datasets were purchased or licensed, with documentation due from 1 January 2026 [5]. A supplier that cannot describe its data in those terms leaves you unable to comply.

In the EU, providers of general-purpose AI models must maintain a copyright policy and publish a summary of training content under Article 53 of the AI Act [6]. As of October 2026, those obligations have applied since 2 August 2025, and the Commission's template for the public summary dates from 24 July 2025 [7]. Write the summary-relevant facts (data categories, collection period, licensed versus other sources) into the sign-off record so the team owning the summary does not have to reconstruct them. The AI training data compliance hub covers the wider regulatory map.

The sign-off memo and handoff records

Counsel's work is only useful if it becomes controls, so close every deal with a short sign-off memo that governance and data engineering can act on. The memo is the record you will reach for when a regulator, plaintiff or acquirer asks how a dataset entered the training corpus.

Illustrative example: invented to show structure; it does not describe an available dataset.

dataset_id: DS-2026-041
license_ref: DLA-2026-017 (executed 2026-09-30)
supplier_role: data owner; records created in own operations
record_classes: [support_tickets, agent_notes]
approved_uses: [fine_tuning, evaluation, rag]
prohibited_uses: [pretraining_redistribution, re_identification, resale]
derivatives: fine-tunes and successor versions permitted; synthetic outputs permitted internal only
model_retention_after_term: weights retained; raw data deleted
data_deletion_date: 2029-09-30
deidentification: token replacement for names, emails, phones, account numbers; 500-record sample, 0 residual direct identifiers
open_risks: [embedded PDF attachments not scanned -> excluded at ingest]
disclosures: [AB 2013 documentation entry, EU training content summary input]
owners: {governance: AI governance lead, pipeline: ML data engineering}
review_date: 2027-09-30

Hand the approved and prohibited uses to the AI governance leads who maintain the use register, and the deletion dates and exclusions to the ML data engineering teams who enforce them in pipelines. For an existing corpus with gaps, a provenance audit of the training corpus is the remediation route.

Red flags that should stop or pause a deal

A short list of conditions should pause a deal until resolved, regardless of commercial pressure. These are the issues that most often convert into deletion, injunction or indemnity disputes.

  • The supplier cannot name the system of record or the collection period.
  • The supplier acts as a processor for its customers' data under a DPA and has no customer consent to license it.
  • "Anonymized" appears without a written method or sample result.
  • Records include call or meeting recordings without evidence of the consent standard for each relevant state.
  • The license grants rights to "the data" but is silent on the trained model after termination.
  • The title warranty is capped at fees paid with no carve-out for third-party IP or privacy claims.

Supplier-side counsel see the mirror image of this review; the legal guide to selling data to AI companies explains what they are told to protect.

How SourceX fits into counsel's review

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, with diligence materials covering source, rights, preparation and allowed use prepared per dataset. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect.

Health records require HIPAA de-identification by Safe Harbor or Expert Determination, and delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Nothing is contracted until a supplier agrees, and a request does not guarantee a match. Counsel can review how SourceX approaches this on its legal framework page, and buyer teams can describe the data they need on the SourceX buyers page. Other team perspectives sit in the data sourcing by buyer team hub.

Bring counsel-ready diligence into your next AI data deal

If your team needs operational data from US companies with diligence materials on source, rights, preparation and allowed use, describe the data rather than the businesses. SourceX runs Find, Assess, Agree, Transact and Manage, with pricing and allowed uses set in a license per deal. Describe your data request on the SourceX buyers page.

Sources

  1. Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
  2. Federal Trade Commission (Office of Technology), "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  3. Federal Trade Commission, "Coming face to face with Rite Aid's allegedly unfair use of facial recognition technology" (2023). https://www.ftc.gov/business-guidance/blog/2023/12/coming-face-face-rite-aids-allegedly-unfair-use-facial-recognition-technology
  4. Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
  5. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  6. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  7. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  8. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  9. International In-house Counsel Journal, "Practical guide on AI training-data risk for in-house counsel". https://iicj.net/paper/3787
  10. Osborne Clarke, "Session 4: AI Licensing" (2024). https://osborneclarke.com/system/files/documents/24/11/21/Session-4---13-Nov---AI-Licensing%28157063266.2%29.pdf
  11. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  12. California Legislature, "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data