Skip to content

AI data market

California AB 2013 explained for companies that license data to AI

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

For companies that license data to AI, California AB 2013 creates no posting duty of their own. The duty sits with developers of generative AI systems offered to Californians, who must publish a high-level summary of their training datasets. That summary can name a supplier and say whether its records hold personal information, so fix the wording in the license.

Key takeaways

  • AB 2013 places its posting duty on developers, not on companies that only supply records.
  • A developer's documentation can describe a supplier's records and may name the supplier.
  • The law does not define high-level precisely, so the license is where a supplier fixes the wording.
  • Operational records raise personal information and company identity questions more than copyright questions.
  • The SourceX Evidence Packet holds most of the facts a developer needs to describe licensed records accurately.

What does California AB 2013 ask of companies that license data?#

California AB 2013, titled Generative Artificial Intelligence: Training Data Transparency, asks nothing directly of companies that only license data. Signed on September 28, 2024 and codified in the Civil Code, it requires developers of generative AI systems or services released on or after January 1, 2022 and made available to Californians to post documentation about their training data on their websites, with a first deadline of January 1, 2026.

Suppliers are affected indirectly. The developer's documentation describes the datasets it used, and a licensed package of support tickets, dispatch records or quality reports is one of those datasets. How the package is described, and whether your company is named, becomes public.

A company that also builds or substantially modifies generative AI systems for public use may be a developer itself. If your company fine-tunes models for its own customers, ask counsel whether that role brings a separate posting duty.

The law's status is worth checking before relying on it. At least one developer challenged it in federal court; a preliminary injunction was denied in March 2026, an appeal was pending as of mid-2026, and commentators differ on how the law would be enforced.

The disclosure items and what suppliers should expect described#

The disclosure items in AB 2013 cover a set of high-level facts about each training dataset, and most of them touch a supplier's records. The table paraphrases the items; confirm the statutory wording with counsel before relying on it.

Read the right-hand column as a forecast, not a rule. Each developer decides its own wording unless the license fixes it.

The disclosure items and what suppliers should expect described
Disclosure item (paraphrased)What it touches in your recordsWhat to expect described
Sources or owners of the datasetsYour company name, or a category such as US logistics providersYour name or a general description, depending on the license
How the datasets serve the system's intended purposeThe permitted use in your licenseA short statement such as improving support or operations agents
Approximate number of data points, which may be expressed as rangesVolume of tickets, jobs, orders or messagesAn approximate range rather than an exact count
Types of data pointsYour record families and fieldsLabels such as support conversations or order exception records
Whether datasets include copyrighted, trademarked or patented material, or are entirely public domainInternal documents and code your company authoredA statement that the data includes material licensed from its owner
Whether datasets were purchased or licensedYour licenseA plain statement that the data was licensed
Whether datasets include personal information, as defined in California's privacy lawWhat remains after de-identificationA yes or no that should match your privacy record
Whether datasets include aggregate consumer informationSummaries of customer behavior, if any were deliveredOften not relevant to record-level operational data; confirm with counsel
Cleaning, processing or other modification by the developer, and its purposeYour own preparation steps, if the developer describes themThe developer's processing, which may also mention de-identification done before delivery
Time period of collection, noting if collection continuesYour archive's date range and any refresh deliveriesA date range, marked ongoing if refreshes continue
Dates the datasets were first used in developmentNothing on your sideThe developer's own dates
Whether synthetic data generation was usedData the developer may generate from your recordsA note on synthetic data, which may mention its sources

How high-level is a high-level summary?#

A high-level summary under AB 2013 is not defined with precision, so each developer decides how much detail to give and practice varies. Some developers describe sources by category; others list named sources or owners. Because the documentation must say whether datasets were purchased or licensed and whether they include personal information, those two answers are the least negotiable.

That flexibility is where suppliers have influence. A license can state the description the developer will use for your records, as long as it does not stop the developer from meeting a legal duty. Agreeing the words in advance is easier than asking for a correction after publication.

Expect documentation to be updated when a developer releases or substantially modifies a system. A license that runs for several years may see your records described more than once, so the agreed description should apply to every version.

Why operational records raise different questions than media archives#

Operational records raise different AB 2013 questions than media archives because the sensitive items are personal information, customer relationships and company identity rather than copyright. Much published commentary on training data disclosure has focused on publishers and creative works, leaving operating companies to work out their own position.

For a logistics firm, a contractor or a manufacturer, being named as a source can reveal which systems it runs and whose work shaped the records. A personal information answer that conflicts with your privacy record can prompt questions from customers. Settle both points while the license is still a draft.

Clauses to settle before signing#

The clauses to settle before signing turn a developer's discretion into agreed terms. They do not prevent compliance; they decide how compliance describes you.

  • An agreed description of your records for any public training data documentation, used in every version.
  • No naming of your company without written consent, except where disclosure is legally required, and then only to the extent required.
  • Volume stated in ranges rather than exact counts.
  • A personal information statement consistent with the privacy record delivered with the package.
  • Advance notice and a review window before documentation that mentions your records is published or changed.
  • Cooperation limited to facts in your documentation set, so you are not asked to certify the developer's own statements.
  • Confidentiality for the documentation schedule itself, so internal detail is not published in full.

Illustrative: a 3PL reviews a developer's draft entry#

Illustrative: a fictional third-party logistics provider licenses warehouse exception records and carrier claim histories from its WMS and TMS to a developer of operations agents offered to the public. Late in negotiation, the developer shares a draft documentation entry that names the 3PL as a source and says the dataset includes personal information.

General counsel compares the draft with the privacy record. Driver names and dock staff IDs appeared in early samples but were removed before delivery, and customer names were replaced with placeholders. The parties agree to describe the package as licensed operational records from a US logistics provider, with personal information removed before delivery, and the 3PL is not named.

Outcome: the license carries the agreed wording, a notice clause for future versions and a cooperation duty limited to facts in the 3PL's own documentation.

How SourceX prepares the facts behind a summary#

SourceX prepares the facts behind a training data summary as part of each package. The SourceX Evidence Packet records provenance, licensing rights, permitted use, the privacy record and release authorization, which together answer most of the items a developer's documentation covers.

The structure follows emerging industry practice: the Data & Trust Alliance's Data Provenance Standards group dataset metadata into Source, Provenance and Use. Within the SourceX five-step transaction, the Rights step records what may be said about the records, and the supplier confirms the naming and description terms at Approval, before the contract is signed.

Frequently asked questions

Does AB 2013 apply to a company that only licenses data?

Generally, the posting duty falls on developers of generative AI systems or services made available to Californians, not on companies that only supply records. A company that also builds or substantially modifies such systems for public use may have its own duty. Counsel should confirm your role for each product and license.

Will the developer publish our records?

No. AB 2013 calls for documentation that describes training datasets at a high level, not publication of the data. Your records stay under the license terms. What becomes public is a description, which is why its wording deserves attention before signing.

Can we stop a developer from naming our company?

You can negotiate it. Many licenses bar naming the supplier without consent, with an exception for disclosure required by law. If a developer's counsel concludes the law requires naming the source, the clause may give way, so narrow that exception to what is actually required and ask for advance notice.

Does AB 2013 change who owns licensed records?

No. It is a transparency law about what developers disclose, not a rule about ownership or permitted use. Ownership and use remain governed by your license, which should state that the company keeps ownership and that rights not expressly granted are reserved.

Do other laws ask for similar summaries?

Yes. Article 53(1)(d) of the EU AI Act requires providers of general-purpose AI models to publish a sufficiently detailed summary of the content used for training, following a template the European Commission published on July 24, 2025. Developers that sell in both markets often send suppliers one questionnaire, so a single documentation set that answers both avoids inconsistent descriptions.

Sources

  • California AB 2013 (Generative Artificial Intelligence: Training Data Transparency), signed September 28, 2024, requires developers of generative AI systems released on or after January 1, 2022 to post training-data documentation on their websites on or before January 1, 2026, stating whether the datasets were purchased or licensed and whether they include copyrighted material or personal information. Source
  • Article 53(1)(d) of the EU AI Act requires providers of general-purpose AI models to make publicly available a sufficiently detailed summary of the content used to train the model, following an AI Office template the European Commission published on July 24, 2025. Source
  • The Data & Trust Alliance's Data Provenance Standards define dataset metadata in three groups, Source, Provenance and Use, which the specification says is needed to enable proper dataset selection for AI model training. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify