Skip to content

Regulation and governance for data buyers

AI Training Data Disclosure Requirements Compared: One Supplier Record for AB 2013, the EU Summary and Colorado

Quick answer

Three regimes now ask model builders to describe their training data: California AB 2013 (a public website summary per dataset, in force since 1 January 2026), the EU training content summary under AI Act Article 53(1)(d) (a public, template-based summary for general-purpose models), and Colorado SB 26-189 (documentation handed to deployers from 1 January 2027). Their fields overlap heavily. Collect one per-dataset record from every supplier, including synthetic-data lineage, and you can fill all three without re-asking.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What each regime requires, as of October 2026

Each regime has a different trigger, audience and granularity, so the same dataset is described three ways. California speaks to the public at dataset level, the EU speaks to the public at model level with source categories, and Colorado speaks to deployers about the system.

California AB 2013. Civil Code Section 3111 requires a developer of a generative AI system or service made available to Californians to post documentation about the data used to train it, including a high-level summary of the datasets [3]. The summary covers sources or owners, how the data furthers the intended purpose, the number of data points (ranges are allowed), the types of data points, whether the data includes copyrighted, trademarked or patented material, whether it was purchased or licensed, whether it includes personal information or aggregate consumer information as defined in the CCPA, any cleaning or processing, the collection period, the dates of first use, and whether synthetic data generation was used [4]. The duty reaches systems released or substantially modified since 1 January 2022, and the posting deadline was 1 January 2026 [3]. Do not confuse it with SB 53, which governs frontier developers' safety frameworks and transparency reports rather than dataset summaries [8].

EU training content summary. Article 53(1)(d) requires providers of general-purpose AI models to publish a sufficiently detailed summary of training content using the AI Office template [6]. The Commission published the explanatory notice and template on 24 July 2025 as a common minimal baseline [1]. The template has three sections: general information about the model and its training data, a list of data sources by type, and data processing aspects; it applies to open-source GPAI providers too [2]. The Article 53 duties have applied since 2 August 2025, with AI Office enforcement powers starting 2 August 2026 for new models [6].

Colorado SB 26-189. Signed on 14 May 2026, it replaced the consumer protections of the 2024 Colorado AI Act [5]. From 1 January 2027, developers of automated decision-making technology that materially influences a consequential decision must give deployers technical documentation covering intended uses, categories of training data and known limitations [5]. It is a business-to-business duty, not a public posting, and it is aimed at decision systems rather than generative models as such.

The deeper treatment of each regime lives on its own page: California AB 2013 training data disclosure, completing the EU summary for licensed and private datasets and Colorado SB 26-189 documentation. The compliance hub maps the rest.

Where the regimes diverge and why it matters for licensed data

The regimes diverge on unit of description, on what they emphasize and on who reads the output. Those differences decide which supplier facts you need at what precision.

DimensionCalifornia AB 2013EU Article 53(1)(d) summaryColorado SB 26-189
TriggerGenerative AI made available to Californians, including substantial modifications [3]Placing a GPAI model on the EU market [6]Covered ADMT that materially influences consequential decisions [5]
AudiencePublic websitePublic, via the AI Office templateDeployers
UnitEach dataset, in a high-level summaryThe model, by source categoryThe system
EmphasisOwnership, counts, IP, personal information, dates, synthetic use [4]Source categories, crawled sources, opt-out compliance, illegal-content handling [2]Categories of data, intended uses, known limitations [5]
Licensed dataAsks whether datasets were purchased or licensed [4]Commercially licensed private data is a source category of its own [2]Folded into categories of training data

Three practical consequences follow. First, California is the most granular at dataset level, so a record that satisfies AB 2013 usually contains what the EU summary needs for licensed and private data. Second, the EU asks about processing at the model level, especially how you honored text-and-data-mining reservations and removed illegal content, which suppliers can only inform for their own preparation steps. Third, Colorado's "known limitations" field is not covered by the other two, so ask suppliers about coverage gaps, label quality and population skew up front.

For licensed data specifically, the EU template treats data obtained from third parties under commercial agreements separately from crawled and public data [2], so the license record itself becomes disclosure evidence. If you fine-tune someone else's model with acquired data, check whether you become a provider or developer in your own right; see fine-tuning provider duties under the EU AI Act and AB 2013.

One per-dataset disclosure record that fills all three

A single record per dataset, with one row per field and a column per regime, removes most re-work. Keep it in your data inventory or a dedicated disclosure register keyed by license ID, and fill it at intake, not at publication time.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldAB 2013EU summaryColoradoSupplier question that fills it
Source owner and supplier legal nameSources or ownersSource category (licensed private data)Categories of training dataWho owns the records and who is licensing them to us?
License ID and acquisition typePurchased or licensedCommercially licensed vs other third-party dataSupporting recordWhich agreement and schedule cover this delivery?
Collection period (start, end, ongoing flag)Time period of collectionLatest collection date at model levelLimitations contextWhen were the earliest and latest records created, and is collection ongoing?
Record count and unit (records, tokens, hours)Number of data points (ranges allowed)Size band by modalityNot requiredHow many records, in what unit, and how was it counted?
Data point types and labelsTypes of data points and labelsModality and nature of contentCategories of training dataWhat fields, file formats and label schema are included?
IP statusCopyright, trademark, patent or public domainCopyright policy inputsNot requiredDoes the data include third-party copyrighted material, and on what basis can you license it?
Personal information status after de-identificationPersonal information and aggregate consumer information under the CCPANo dedicated template fieldLimitations contextWhich identifiers were removed or replaced, by what method, and what remains?
Processing stepsCleaning, processing or modification and purposeData processing aspectsKnown limitationsWhat filtering, deduplication, redaction or normalization was applied?
Synthetic share and generatorWhether synthetic data generation was usedSynthetic data categoryCategories of training dataIs any content synthetic; if so, what share, which generator model and which seed data?
First-use dateDates first used in developmentNot requiredNot requiredFilled internally by the buyer
Known limitationsNot requiredNot requiredKnown limitationsWhat populations, periods or cases are missing or under-represented?

Two fields are always buyer-side: first-use date and how the dataset furthers the system's purpose. Log them at the training run, since nobody else can reconstruct them later. Model lineage and permitted uses belong in a separate use register; the provenance hub covers that, and the licensing hub covers the license terms that feed the record.

Disclosing synthetic data and its seed licenses

AB 2013 asks whether the system used, or continuously uses, synthetic data generation, and allows a description of its purpose [4]. The EU template lists synthetic data among its data source categories alongside public, licensed, crawled and user data; check the template file itself for the exact fields before you publish [1][2].

Synthetic data generated from licensed seed data does not escape the seed license. If the license limits use to a named purpose or model family, synthetic output typically inherits that limit, and a generator model's own output terms can add restrictions. Record three things for every synthetic component: the generator model and version, the terms under which its outputs may be used for training, and the license IDs of the seed data. The guide on combining licensed and synthetic data and the synthetic data glossary entry go further.

A common failure mode is a supplier that ships "augmented" data, such as paraphrased support tickets or LLM-generated summaries of documents, without flagging it as synthetic. Ask directly, and require a per-file or per-record flag rather than a dataset-level yes or no.

A supplier request that answers every regime

One request at the Assess stage avoids a second round of questions during publication. Send it with the record template above and ask for answers per dataset, not per supplier.

Illustrative example: invented to show structure; it does not describe an available dataset.

Subject: Disclosure facts for dataset <license_id>

Please provide, per dataset in the schedule:
1. Owner of the records and licensing entity (legal names)
2. Collection window: earliest and latest record dates; ongoing yes/no
3. Count and unit (records / tokens / audio hours); counting method
4. Field list, file formats, label schema and labeling process
5. IP status: third-party copyrighted, trademarked or patented content; basis for licensing
6. Personal information: categories present before preparation; identifiers removed or
   replaced; method used; sample check performed
7. Processing applied: filtering, dedup, redaction, normalization, with tool names
8. Synthetic content: yes/no per file; share; generator model and version;
   generator output terms; seed data license IDs
9. Known limitations: coverage gaps, skew, label error rates if measured
10. Any opt-out or rights reservation you honored when assembling the data

Treat answers as evidence: attach them to the license record and version them. Item 10 informs your EU copyright policy work under the GPAI Code of Practice, which is voluntary but widely used as a reference [7]. For broader diligence, pair this request with the training data due diligence checklist and the EU summary guide for buyers. If you are sourcing new licensed data, you can describe the dataset you need to SourceX and use this request as part of the Assess stage.

Common mistakes in a training data disclosure register

Most disclosure errors come from facts that were never captured, not from misread statutes. The recurring ones:

  • Counting in different units. A supplier reports documents, your pipeline counts tokens after chunking, and the AB 2013 range no longer matches the EU size band. Store the supplier's unit and your post-processing unit side by side.
  • Using license dates as collection dates. The license signature date is not when the data was created. AB 2013 asks for the collection period [4].
  • Losing personal information status after processing. Record the status after de-identification and the method used, because "contains personal information" at delivery and at training can differ.
  • Treating Colorado as a copy of AB 2013. Colorado's documentation goes to deployers and asks for limitations [5], which a public summary does not.
  • Forgetting refreshes. An ongoing feed changes counts, dates and sometimes synthetic share. Version the record per delivery.

Sourcing licensed data with documented source and rights

SourceX sources operational datasets from US companies on request and manages the licensing process; every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery, with diligence materials on source, rights, preparation and allowed use prepared per dataset. Personal details are removed or replaced before delivery and the method is recorded. Describe the data you need at SourceX for AI data buyers.

Sources

  1. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  2. WilmerHale, "European Commission Releases Mandatory Template for Public Disclosure of AI Training Data" (2025). https://wilmerhale.com/en/insights/blogs/wilmerhale-privacy-and-cybersecurity-law/european-commission-releases-mandatory-template-for-public-disclosure-of-ai-training-data
  3. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  4. Conventus Law, "US: California's AB 2013 Requires Generative AI Data Disclosure By January 1, 2026". https://conventuslaw.com/report/us-californias-ab-2013-requires-generative-ai-data-disclosure-by-january-1-2026/
  5. Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
  6. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  7. European Commission (AI Office), "The General-Purpose AI Code of Practice" (2025). https://digital-strategy.ec.europa.eu/en/policies/gpai-code-practice
  8. Morrison & Foerster LLP, "At the Frontier - California Enacts AI Safety and Transparency Regulation TFAIA (SB 53)" (2025). https://www.mofo.com/pdf/resources/insights/251001-california-enacts-ai-safety-transparency-regulation-tfaia-sb-53

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data