Data licensing for AI training
Naming your data sources: confidentiality clauses vs AI transparency duties
Quick answer
A standard confidentiality clause can make it a breach to name a licensor or describe its data, while California AB 2013, EU AI Act Article 53 and your own model cards may require exactly that. The fix is contractual, not editorial. Negotiate a disclosure carve-out for legally required documentation, agree in advance on the description level the licensor accepts (category, sector, size, time range), and keep regulatory disclosure separate from publicity rights such as press releases and logos.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Where confidentiality and transparency collide
The collision happens because several disclosure duties ask about the source of training data, and data licenses often treat the deal itself as confidential. As of October 2026, the main pressure points for a buyer that trains models are these:
- California AB 2013 (Civil Code 3110 et seq.). Developers of generative AI systems made available to Californians had to post training-data documentation by January 1, 2026. The required high-level summary includes the sources or owners of the datasets and whether the datasets were purchased or licensed by the developer [1].
- EU AI Act Article 53. Providers of general-purpose AI models must keep technical documentation, give information to downstream providers, maintain a copyright policy and publish a sufficiently detailed summary of training content; these duties have applied since 2 August 2025 [2]. The Commission's template for that public summary, dated 24 July 2025, sets a common minimum baseline [3], and the voluntary GPAI Code of Practice covers transparency and copyright [4].
- Colorado SB26-189. From January 1, 2027, developers of automated decision-making technology used for consequential decisions must give deployers documentation that includes categories of training data [5].
- Voluntary documentation. Data Cards and Hugging Face dataset cards expect upstream sources and license information to be described [9][10], and Croissant-RAI makes that metadata machine-readable [11].
None of these duties is waived because a contract says the data is confidential. If your license forbids describing the source, you are choosing between breach of contract and an incomplete or inaccurate filing.
Why licensors resist being named
Licensors resist disclosure because naming them can reveal a commercial relationship, a data asset competitors did not know existed, or the fact that customer records left the building at all. Commentators on AB 2013 have focused on the risk that training-data documentation exposes trade secrets [6], and xAI challenged the statute in federal court on trade secret and compelled-speech grounds [7]. The status of that challenge can change, so check it before relying on any outcome.
There is also a structural reason. Under the Defend Trade Secrets Act definition, information keeps trade secret status only if the owner takes reasonable measures to keep it secret [8]. A licensor that lets buyers describe its data freely may worry it is weakening those measures, so it defaults to a blanket confidentiality clause. Your job is to show that a bounded, pre-approved description does not do that.
The disclosure carve-out: what the clause must cover
A workable carve-out permits disclosure that is required by law or regulation, or that is made in model documentation, at a description level agreed in advance. Most template NDAs already carve out disclosures compelled by law or court order, but that carve-out is usually too narrow for AI documentation, for three reasons.
First, "compelled by law" often triggers notice and cooperation steps designed for subpoenas, not for routine public filings such as the AB 2013 web posting. Second, voluntary documentation (model cards, dataset cards, system cards, responses to enterprise customer questionnaires) is not compelled at all. Third, Article 53 information that goes to the AI Office or downstream providers is a different audience from the public summary, so the clause should treat regulator-only and public disclosures separately [2].
Draft the carve-out around four elements:
- Covered documents. Name them: public training-content summaries, statutory website postings, technical documentation for regulators, downstream-provider documentation, model and system cards, and dataset cards.
- Approved description. Attach a schedule with the exact wording or the permitted fields.
- Regulator-only channel. Allow fuller detail, including the licensor's name, to regulators and auditors under their own confidentiality rules.
- Change control. Require licensor approval only for disclosures beyond the schedule, with a fixed review window and deemed approval if it lapses.
Illustrative example: invented to show structure; it does not describe an available dataset.
X.4 Permitted Transparency Disclosures.
Notwithstanding Section X.1 (Confidentiality), Licensee may disclose:
(a) the Approved Description in Schedule D in any public training-data
summary, statutory website posting, model card, system card or
dataset card relating to a model trained in whole or part on the
Licensed Data;
(b) such further information, including Licensor's identity, as a
competent regulator, auditor or downstream provider requires under
applicable law, provided the recipient is bound by confidentiality
obligations by law or contract;
(c) any other description with Licensor's prior written approval,
which Licensor will not unreasonably withhold.
X.5 This Section does not grant any right to use Licensor's name,
marks or logos in marketing, press releases or customer lists,
which remain governed by Section Y (Publicity).
Agreeing the description level before signature
The description level is the most negotiable part of the clause, and it should be fixed in the license, not argued about when the filing is due. Build a schedule that maps each documentation field to the wording both sides accept.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Documentation field | Licensor named | Sector-level (common compromise) | Minimal |
|---|---|---|---|
| Source or owner | "Acme Logistics Inc." | "A US freight logistics company" | "A commercial data licensor" |
| Acquisition route | Licensed under agreement dated X | Licensed from a third party | Licensed |
| Data type | Dispatch tickets and driver chat logs | Operational support tickets and messages | Text records |
| Volume | 4.2M tickets | 1M to 10M records | Not stated |
| Time range | Jan 2019 to Dec 2024 | 2019 to 2024 | Not stated |
| Personal information | Pseudonymized; method on file | Contains de-identified personal information | Contains personal information |
| IP status | Licensor-owned, copyright asserted | Contains copyrighted material, licensed | Licensed content |
Check the minimal column against each regime before accepting it. AB 2013 asks about sources or owners and whether data was purchased or licensed [1], so "a commercial data licensor" may be thin; counsel should decide whether a sector-level description meets "high-level summary." For the EU public summary, the exact fields are in the Commission's template documents, so map them to your schedule directly [3]. For the personal-information row, align the wording with the de-identification evidence you received; see the de-identification evidence package checklist.
Keep publicity rights out of the transparency clause
Publicity and transparency are different rights and should sit in different clauses. A publicity clause controls marketing uses: press releases, logo walls, case studies and sales decks. A transparency clause controls legally required or documentation-driven descriptions of training data.
Merging them causes two failure modes. If the publicity clause is strict ("no use of Licensor's name without consent") and the transparency carve-out cross-references it, your AB 2013 posting is blocked by a marketing restriction. If the transparency carve-out is generous and reads as a general license to name the licensor, sales teams will treat it as permission to announce the partnership. Keep Section X.5 above, or an equivalent, so each right stays in its lane.
Model cards, dataset cards and internal documentation
Model and dataset cards are voluntary, so they need explicit coverage in the carve-out or they default to the confidentiality clause. The Data Cards framework expects upstream sources, collection methods and intended use to be described [9], and a Hugging Face dataset card publishes license, size and other metadata in a public README [10]. If you publish a card for a fine-tuned model, the training-data section must stay within the Approved Description.
Internal documentation needs a rule too. Your governance file should hold the licensor name, agreement reference, record counts and the full provenance trail, marked confidential, while the public card points only to the schedule wording. Croissant-RAI metadata can carry both if you control which fields are exported [11]. For how dataset cards work with licensed enterprise data, see dataset cards for licensed enterprise data.
Pre-signature checklist for governance leads
Use this list before the license is executed; it is generic buyer practice, not a description of any vendor's terms.
Illustrative example: invented to show structure; it does not describe an available dataset.
- List every model or system the data may train and every regime that applies (AB 2013, EU Article 53, Colorado SB26-189, customer contract requirements).
- Confirm the confidentiality clause has a transparency carve-out broader than "compelled by law."
- Attach a Schedule D Approved Description covering source, acquisition route, data type, volume, time range, personal information and IP status.
- Confirm a regulator-only channel for fuller detail.
- Separate publicity (name, logo, press) from transparency.
- Check that the confidentiality term does not outlast your duty to keep the documentation current.
- Check how the carve-out interacts with deletion and return obligations, since documentation about a model often outlives the data.
- Ask the supplier for the diligence materials you will need to support the description (source, rights, preparation, allowed use).
The rights grant itself is covered in the AI training rights grant clause, and the full negotiation list is in the data license negotiation checklist. For what to request from suppliers to build an EU summary, see EU AI Act training data summaries: what buyers need from suppliers.
How a sourcing intermediary changes the disclosure question
Working through an intermediary does not remove the need for an Approved Description, but it changes who negotiates it. SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements; every release is approved by the supplying company, and each dataset is delivered under a license defining records, uses, term and delivery. Buyers describe the data they need, not the businesses, so raise documentation and disclosure requirements at the start of the request on the SourceX buyer page. Pricing and allowed uses are set in the license at the Agree step, and nothing is contracted until a supplier agrees. For keeping your own roadmap confidential during sourcing, see confidential data sourcing, and for the wider cluster, the AI training data licensing guide.
Sourcing licensed data you can document
SourceX sources operational datasets from US companies on request, with every dataset rights-reviewed and delivered under a license that defines records, uses, term and delivery. A request does not guarantee a match, and terms are agreed per deal. Describe the data you need and your documentation requirements at sourcex.si/buyers.
Frequently asked questions
Can we name a licensor in a model card if the NDA is silent?
Not safely. A silent NDA usually means the general confidentiality clause governs, and the licensor's identity and deal terms are often defined as confidential information. Get written approval or amend the license to add a carve-out.
Does an EU training-content summary require naming every licensor?
The summary uses the Commission template as a minimal baseline [3], and the exact level of detail for licensed data is set in that template and its explanatory notice. Read those documents against your Approved Description; do not rely on secondary summaries.
Is a regulator-only disclosure still a breach of confidentiality?
It can be if the license does not permit it. Article 53 contemplates information going to the AI Office and downstream providers [2], so draft the regulator-only channel explicitly instead of relying on a "compelled by law" carve-out.
Sources
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- European Commission (AI Office), "The General-Purpose AI Code of Practice" (2025). https://digital-strategy.ec.europa.eu/en/policies/gpai-code-practice
- Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
- Goodwin Procter, "California's AB 2013 Takes Effect: Navigating AI Training Data Transparency and Trade Secret Risk" (2026). https://www.goodwinlaw.com/en/insights/publications/2026/01/alerts-otherindustries-californias-ab-2013-takes-effect
- Institute for Law & AI, "xAI's Challenge to California's AI Training Data Transparency Law (AB2013)". https://law-ai.org/xais-challenge-to-californias-ai-training-data-transparency-law-ab2013/
- U.S. Government Publishing Office (govinfo), "18 U.S.C. 1839 Definitions (United States Code, 2021 edition)" (2021). https://www.govinfo.gov/content/pkg/USCODE-2021-title18/html/USCODE-2021-title18-partI-chap90-sec1839.htm
- Pushkarna, Zaldivar, Kjartansson (Google Research), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
- Jain et al. (MLCommons Croissant RAI task force), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.