Regulation and governance for data buyers
California AB 2013: Training Data Disclosures and the Records to Collect From Data Suppliers
Quick answer
California AB 2013 requires any developer of a generative AI system made available to Californians to post, on its website, a high-level summary of the datasets used to train it, covering sources and owners, volume, data types, IP status, whether data was purchased or licensed, personal information, processing, collection periods, first-use dates and synthetic data [1]. It applies from January 1, 2026 [1]. For licensed data, collect each fact from the supplier before release, under terms that permit public disclosure.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What AB 2013 requires and who it covers
AB 2013 is a website-posting duty, not a filing: the developer publishes documentation about training data on or before January 1, 2026, and again before each new release or substantial modification [1]. The bill was approved and chaptered on September 28, 2024 as Chapter 817, Statutes of 2024, adding Title 15.2 to the Civil Code starting at Section 3110 [2]. It reaches systems released on or after January 1, 2022, so legacy models still on the market are in scope [1].
"Developer" is broad. It includes anyone who designs, codes, produces or substantially modifies a generative AI system or service for public use [3]. In practice, a team that fine-tunes an open-weight model on licensed records and ships it in a public product should assume it is a developer for that release. The statute treats training to include testing, validating and fine-tuning, so eval sets and fine-tuning corpora belong in the summary too [1]. For the fine-tuning boundary in more depth, see developer duties after fine-tuning under AB 2013 and the EU AI Act.
The twelve summary items, and the ones licensed data triggers
Section 3111 lists twelve items, and at least four depend on facts only a data supplier holds [1]. The table maps each item to the supplier record that answers it.
Illustrative example: invented to show structure; it does not describe an available dataset.
| AB 2013 item [1] | What the developer must state | Supplier record to collect |
|---|---|---|
| Sources or owners of the datasets | Who the data came from | Legal entity name of the rights holder, or an approved generic descriptor if the license restricts naming |
| How the datasets further the system's purpose | Why this data was used | Buyer-side: intended-use statement from the data request |
| Number of data points | Volume; ranges or estimates allowed for dynamic data | Record counts per delivery, unit definition (ticket, message, document, page) |
| Types of data points | Labels for labeled data, general characteristics for unlabeled | Schema or data dictionary, label taxonomy, modality and language |
| Copyright, trademark or patent protection, or public domain | IP status | Ownership representation and description of third-party content embedded in records |
| Whether purchased or licensed by the developer | Acquisition mode | License or order-form identifier and effective date |
| Personal information (Civ. Code 1798.140) | Yes or no | De-identification method, residual identifier fields, sample-check result |
| Aggregate consumer information (1798.140) | Yes or no | Whether any aggregate or cohort statistics are included |
| Cleaning, processing or modification, and its purpose | Pipeline description | Processing log: redaction, deduplication, normalization, filtering |
| Collection time period, and whether ongoing | Date range | Earliest and latest record timestamps; whether the feed is recurring |
| Dates first used in development | First-use date | Buyer-side: ingest date from your training run registry |
| Use of synthetic data generation | Yes or no, with purpose | Whether the supplier generated or augmented any records |
Two items deserve care. The personal information item is tied to the CCPA definition, so data that meets the CCPA "deidentified" conditions in Section 1798.140 is treated differently from merely pseudonymized data [6]. Ask suppliers for the method, not just a "no PII" assertion, and read the obligations a buyer inherits with CCPA deidentified data.
Reading "high-level" when a license limits what you can say
AB 2013 does not define "high-level," and it has no trade-secret exception, so the developer decides granularity and carries the risk [4]. The statute has no dedicated penalty provision [1], and commentators expect enforcement through California's Unfair Competition Law. That leaves two failure modes: a summary so vague it misstates the data, or one detailed enough to breach a confidentiality clause in a data license.
Resolve the tension in the contract, before data moves. Agree on a disclosure descriptor with each supplier, such as "licensed customer support transcripts from a US software company, 2019 to 2024," and record it as an approved public statement. Agree also that volume may be published as a range. Because the statute offers no confidentiality shield, decide the wording in advance rather than at release [4]. For clause design, see disclosure requirements compared across AB 2013, the EU summary and Colorado.
Exemptions and their limits
The exemptions are narrow and system-based, not dataset-based [5]. Section 3111 excludes a generative AI system whose sole purpose is security and integrity, one used solely for operating aircraft in national airspace, and one developed for national security, military or defense purposes and made available only to a federal entity [1]. A general assistant that also does security tasks does not qualify. Nothing exempts licensed, proprietary or confidential datasets from the summary.
Legal challenge and status as of October 2026
As of October 2026, AB 2013 is in force, and a federal constitutional challenge filed by xAI in December 2025 has been reported in the press, arguing trade-secret takings, compelled speech and vagueness of terms such as "dataset" and "data point." Check the docket status with counsel before relying on any outcome; this page makes no prediction.
Do not plan on the law disappearing. Build records that let you comply now, and keep the descriptor approach above so a narrower or broader reading later needs only a rewrite of the posted summary, not a fresh supplier inquiry. Watch for Attorney General guidance as well; no formal regulations are attached to the statute [1].
Supplier record checklist before release
Collect these fields per dataset at contracting and freeze them at the training-run ingest date. This is a buyer-side artifact; every supplier will hold different evidence.
Illustrative example: invented to show structure; it does not describe an available dataset.
dataset_id: DS-2026-031
license_ref: LIC-0412 (effective 2026-03-02)
acquisition: licensed # purchased | licensed | first-party | public
rights_holder: "US B2B software company" # approved public descriptor
named_disclosure_allowed: false
modality: text; language: en-US
unit: support ticket thread
record_count: 1,240,000 # publish as "1M-2M"
labels: [issue_category, resolution_code, csat_bucket]
ip_status: supplier-owned; third-party content limited to customer-authored text
personal_info: deidentified # method: NER redaction + token replacement
deid_method_ref: DEID-v3, sample QA 2,000 records
aggregate_consumer_info: false
processing: [redaction, dedup (MinHash), language filter, HTML strip]
collection_period: 2019-01 to 2025-12; ongoing: false
first_used: 2026-04-18 (run sft-2026-04)
synthetic_generation: false
Map each field to its AB 2013 item, then check four common gaps: no first-use date recorded because ingest logs were rotated; volume counted in tokens by one team and records by another; a supplier that augmented records with an LLM without saying so; and personal-information status claimed without a documented method.
Reusing the same record for the EU summary and other laws
One supplier record can feed AB 2013 and the EU training-content summary, but the two differ in scope and depth. The European Commission's template for the Article 53(1)(d) public summary, published 24 July 2025, applies to general-purpose AI model providers and asks for structured information about data sources [7]. AB 2013 covers any public generative AI system, including fine-tuned applications [1]. Keep one internal schema and render two outputs. For the EU side, see completing the EU summary template for licensed and private datasets, and for the US landscape, US state AI laws that reach training data. Retention matters too: AB 2013 sets no retention period, but you will want first-use evidence to survive the end of a license, which is covered in training data retention requirements.
How SourceX fits AB 2013 record-keeping
SourceX sources operational datasets from US companies on request and manages licensing and ongoing purchases. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, and personal details are removed or replaced before delivery with the method recorded and a sample checked. No method is perfect. Those materials map onto the supplier columns above; you can describe what you need on the SourceX buyer page. For US privacy context, see our United States law overview and CCPA and licensing business data. More compliance guides are in the compliance hub and the AI data hub.
Licensed training data for an AB 2013 summary
SourceX sources operational datasets from US companies on request, and every release is approved by the supplying company. Each dataset arrives rights-reviewed, under a license that defines records, uses, term and delivery, with per-dataset diligence materials you can map to your disclosure. Describe the data you need at sourcex.si/buyers.
Frequently asked questions
Does AB 2013 apply to models that are only used internally?
The duty attaches to systems or services made available to Californians, so a purely internal tool with no public access falls outside it [1]. Customer-facing features built on the same model do not.
Do eval and test sets need to be disclosed?
The statute's definition of training includes testing, validating and fine-tuning, so held-out evaluation sets used in development belong in the summary [1]. See also regulatory requirements for AI validation and test data.
Can a data license forbid naming the supplier?
It can, and AB 2013 has no trade-secret carve-out to fall back on [4]. Negotiate an approved generic descriptor that accurately states the source type, so the summary is truthful without naming the company.
Sources
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- California Legislature, "Bill History - AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billHistoryClient.xhtml?bill_id=202320240AB2013
- Conventus Law, "US: California's AB 2013 Requires Generative AI Data Disclosure By January 1, 2026". https://conventuslaw.com/report/us-californias-ab-2013-requires-generative-ai-data-disclosure-by-january-1-2026/
- Perkins Coie, "AB 2013: New California AI Law Mandates Disclosure of GenAI Training Data". https://perkinscoie.com/insights/update/ab-2013-new-california-ai-law-mandates-disclosure-genai-training-data
- Securiti, "Assembly Bill 2013: Generative Artificial Intelligence Training Data Transparency". https://securiti.ai/blog/assembly-bill-2013-generative-artificial-intelligence-training-data-transparency/
- California Legislature, "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.