Skip to content

Due diligence checklist for licensed AI training data

Before licensing AI training data, confirm that the licensor owns or controls the data and may license it for your use; that contracts, notices and consents allow that use; that personal, special-category, confidential, open-source and export-controlled content is removed or cleared; that delivery is secure and documented; and that a representative sample passes your quality and contamination checks. Record the answers as representations and warranties in the license.

Key takeaways

  • Custody is not ownership, so confirm who owns the data and whose authorization is needed to license it.
  • The contracts, notices and consents in force when the data was collected limit what it can be licensed for now.
  • Verify de-identification on a sample, including free text, attachments and metadata, rather than relying on a description of the method.
  • Health, biometric and children's data, privileged material, copyleft code and export-controlled data each need a specific check or exclusion.
  • Write the answers into the license as warranties and keep a provenance record for every delivery.

Run these checks before you sign, against the licensor's written answers and a sample prepared like the full delivery. An unchecked box means a fix, an exclusion or a dropped candidate. Then write the answers into the license; AI data license terms explained covers the clauses.

Rights and permissions

Ownership and chain of title

Chain of title is the documented path by which the licensor came to hold the rights it grants; custody alone does not establish it.

  • The licensor is the entity that owns or controls the data, not only one that hosts or processes it.
  • Where the holder handles data for clients, as contact centers, agencies and law firms do, each client has authorized licensing in writing.
  • Data that came with an acquisition or merger is covered by the deal documents.
  • Employee and contractor work product is assigned to the licensor.
  • No existing exclusive license, security interest or dispute covers the same data or field of use.

Third-party and customer contract restrictions

Contracts the holder signed with customers, vendors and platforms can restrict what it may do with data, including data it owns.

  • Customer agreements and data processing agreements permit the use, or the affected customers' data is excluded.
  • Terms of the platforms the data was exported from do not prohibit the use.
  • Third-party content inside the data, such as licensed images, market data, standards and vendor manuals, is excluded or separately licensed.

What people were told when their data was collected limits what it can be licensed for now.

  • Privacy notices in force at collection, and the lawful basis relied on, support disclosure for AI development.
  • Call recordings met consent rules where each party was located, including all-party-consent jurisdictions, and the recording disclosure covers the new use.
  • Employees were notified, and works councils consulted, where local law requires it.
  • Opt-outs, objections and deletion requests already received are applied to the delivery.
  • For new recordings, workers gave written consent covering AI training, licensing and retention, and bystanders are kept out of frame or blurred.

Personal and sensitive data

Personal data and de-identification

De-identification should happen before data leaves the holder, and you should verify the result on the sample.

  • A written spec covers structured fields, free text, attachments, images, audio and file metadata.
  • The spec chooses redaction or consistent pseudonyms, and the pseudonym key stays with the holder or is destroyed.
  • Quasi-identifiers, such as rare job titles, small locations and exact dates, are assessed in combination.
  • The legal standard is named where one applies, such as HIPAA de-identification or the CCPA definition of deidentified.
  • The sample passes automated and manual scans for names, contact details, account and card numbers, and government IDs.
  • The license prohibits re-identification and linking with other data to identify people.

Special categories: health, biometrics and children

Health, biometric and children's data carry stricter rules in many jurisdictions; exclude them unless you need them.

  • Health data from a HIPAA covered entity or business associate is de-identified under 45 CFR 164.514(b), by Safe Harbor or Expert Determination, because protected health information generally may not be sold without the individual's authorization.
  • Health details outside HIPAA, such as injuries in property and casualty claims, are handled as special-category data under the GDPR or sensitive data under state privacy laws, where those apply.
  • Voice and face data are assessed: under the GDPR they become special-category biometric data when processed to uniquely identify a person.
  • No biometric identifiers, such as voiceprints or face geometry scans, are delivered or later derived. Illinois's Biometric Information Privacy Act requires a written release to collect them and bars selling or otherwise profiting from them.
  • Children's data is excluded or cleared. The FTC's amended COPPA Rule requires covered operators to get separate verifiable parental consent before disclosing children's personal information to third parties, unless the disclosure is integral to the service.

Confidentiality and privilege

Business records often hold information the holder must keep confidential for others, some of it privileged.

  • Client, counterparty and supplier confidential information, including trade secrets, is excluded or cleared under the governing NDAs and engagement terms.
  • Attorney-client communications and attorney work product are excluded, since disclosing them to third parties can waive privilege.
  • Material under court protective orders or regulatory confidentiality rules is excluded.
  • Passwords, API keys, tokens and private keys are removed from tickets, logs and documents.

Code and technical data

Open-source license contamination in code

Proprietary repositories routinely contain open-source and third-party code under its own licenses, which the company's license cannot override.

  • A license scan lists the licenses found, ideally as SPDX identifiers, including vendored dependencies.
  • Copyleft code, such as GPL, AGPL, LGPL or MPL, is excluded or tagged; how copyleft terms apply to model training is unsettled.
  • Notice requirements of permissive licenses such as MIT, BSD and Apache-2.0 are recorded.
  • Secrets scanning covers the full git history, not only the current tree, and commit identities are pseudonymized.

Export controls on technical data

Engineering data such as CAD models, schematics, source code and manufacturing know-how can be export-controlled, and showing it to a foreign person can count as an export.

  • The holder has classified the data under the regimes that apply, such as the US ITAR and EAR or the EU dual-use regulation.
  • You know which foreign persons will access it: under US rules, releasing controlled technical data, technology or source code to a foreign person inside the United States is a deemed export (22 CFR 120.50, 15 CFR 734.13).
  • Storage locations, cloud regions and contractor access are reviewed.
  • Controlled data is excluded unless you hold the authorizations and controls to receive it.

Delivery, documentation and quality

Security of delivery

Agree how data moves, who can access it and how it will be destroyed before transfer.

  • Transfer starts only after the agreement is executed, the holder approves release and named recipients are authorized.
  • Data is encrypted in transit and at rest and lands in an access-controlled environment.
  • A manifest lists every file with a checksum, such as SHA-256, to verify completeness.
  • Access is limited to named people and logged; onward sharing needs permission.
  • Incident notification and end-of-term destruction are agreed, for example by reference to NIST SP 800-88 media sanitization guidelines.

Provenance documentation

Provenance documentation records where each delivered record came from and what was done to it; it is your evidence if a regulator or court asks.

  • Source systems, export dates, time range and filters are recorded.
  • The rights basis is summarized: owner, authorizations and notices relied on.
  • A processing log covers the de-identification method and version, normalization, deduplication and exclusions, with reasons.
  • A data dictionary describes fields, label definitions and known gaps, such as rubric changes over time.
  • Each delivery has an ID tied to the license, so every training run can be traced to its source.

The EU AI Act expects such records: data governance for high-risk systems must cover data collection and origin, and general-purpose AI model providers must publish a summary of training content.

Sample QA

Sample QA tests a representative sample against acceptance criteria you wrote before seeing it.

  • The sample is random or stratified across the full scope and prepared like the delivery.
  • Records parse against the schema; encodings and timestamps are consistent.
  • Pseudonymous IDs join across tables and over time.
  • Label and outcome coverage meets the spec.
  • Duplicates and templated content, such as macros and auto-replies, are measured.
  • The sample is licensed for evaluation only, with deletion if no license follows.

Contamination of evaluation data

An eval set is contaminated when its items or close variants appeared in a model's training data, which inflates scores; licensed data avoids this only while it stays unpublished and held out.

  • The records were never published, for example in public repositories, forums, help centers or court filings.
  • A near-duplicate search against public corpora and benchmarks finds no meaningful overlap.
  • You know whether the same records are, or may be, licensed to others who could train on them.
  • Splits keep each customer, case or repository on one side, and eval items postdate training data where possible.
  • Eval data stays out of every training run and is not sent to external model APIs that may retain inputs.

Where SourceX fits

SourceX checks fit, rights and quality during qualification, before a license is signed. Partners confirm their licensing rights before delivery, sensitive data is de-identified to requirements agreed in scope, and delivery requires the partner's approval, an executed agreement and authorized buyer access. Treat that as input to your own review, not a substitute for it. Supply is not guaranteed. Send a data request with your spec and diligence requirements.

Questions

Who does due diligence when data is licensed through an intermediary?

Both sides do. SourceX, as the enterprise data transaction layer for AI, checks fit, rights and quality before offering a dataset and coordinates de-identification with the data owner. You should still review the answers with your own counsel against your intended use, because the license allocates the remaining risk. Confirm which party gives each warranty, and whether the data owner signs the license or only the intermediary does.

What documents should a data licensor provide before signing?

At minimum: a manifest describing sources, time range, volume and record contents; a written de-identification spec; a summary of the rights basis, including client authorizations where the holder processed data for others; a data dictionary with known gaps; and a representative sample prepared the same way as the full delivery. For code, add a license scan; for engineering data, an export classification.

Does de-identified data fall outside privacy law?

It depends on the standard and on who holds what. Health information de-identified under HIPAA's Safe Harbor or Expert Determination method is no longer protected health information. Under the GDPR, anonymous data is outside the regulation, but pseudonymized data remains personal data for anyone who can reasonably re-identify it. California's CCPA treats data as deidentified only if conditions are met, including contractual commitments from recipients not to re-identify it.

How do I keep a licensed evaluation set free of contamination?

Use records that were never published, check them for near-duplicates in public corpora and benchmarks, and keep them out of every training run. Split by customer, case or repository rather than by row, so related records stay on one side. Restrict access, avoid sending items to external model APIs that may retain inputs, and consider exclusivity over the eval snapshot so other buyers cannot train on the same records.

Ready to source data?

Send the domain, modality, volume, format, timeline and permitted use you need. SourceX will match it against partner businesses.

Updated 3 October 2026.

See if you qualify