Skip to content

Regulation and governance for data buyers

US State AI Laws That Reach Training Data: What California and Colorado Require and What to Track

Quick answer

As of October 2026, two enacted state laws stand out for putting duties directly on training data. California AB 2013 has required developers of public generative AI systems to post a high-level training-data summary since January 1, 2026, including after substantial modifications such as fine-tuning. Colorado SB26-189 requires developers of covered automated decision-making technology (ADMT) to give deployers documentation that includes categories of training data from January 1, 2027. Most other state AI laws govern how AI is used, not what it was trained on.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Which state laws actually reach training data

Only laws that name training data as the subject of a record belong on a training-data tracker, and in October 2026 that list is short. The test is simple: does the statute require the developer to document, disclose or hand over information about the data used to train, test or modify the model? If it only regulates notices to people affected by an AI decision, bias audits or safety frameworks, it is an adjacent obligation that may still draw on your data records, but it does not create a training-data record of its own.

The table below is the working map for a US developer that acquires third-party data. Each row is dated; re-check before relying on it.

StateLawStatus and key date (as of October 2026)Reaches training data?Record it needsDetailed page
CaliforniaAB 2013, Civil Code §§ 3110–3111Chaptered September 28, 2024; posting required by January 1, 2026 [1]Yes, directlyPublic website summary of datasets used to train the system [1][2]California AB 2013 training data disclosure
ColoradoSB26-189 (ADMT)Signed May 14, 2026; core duties from January 1, 2027; AG rulemaking open [3][4]Yes, as part of developer-to-deployer documentationCategories of training data, intended uses, known limitations, given to deployers [3]Colorado SB 26-189 documentation
ColoradoSB 24-205 (Colorado AI Act)Replaced by SB26-189 in May 2026 [3]No longer the operative textNone; remove from trackersn/a
CaliforniaSB 53 (TFAIA)Signed September 29, 2025; core obligations from January 1, 2026 [5]Indirectly at mostFrontier safety framework and transparency reports for models above 10^26 operations [5]n/a
California, IllinoisEmployment AI rules and disclosure lawIn effect (Illinois HB 3773 effective Jan 1, 2026) [6]No; governs AI use in employment decisionsNotices, anti-discrimination recordsn/a

If you are building one document set to satisfy several regimes, the comparison of AI training data disclosure requirements maps AB 2013, Colorado and the EU summary onto one supplier record.

California AB 2013: a public summary that follows the model, not the vendor

AB 2013 requires the developer of a generative AI system or service made available to Californians to post documentation about the data used to train it on the developer's website [1]. The duty attaches to systems released on or after January 1, 2022, and it was due by January 1, 2026 [1][2]. It also applies when a developer substantially modifies a system, which commentators read to cover fine-tuning and retraining [2].

Section 3111 lists the contents of the "high-level summary" [1]. For a buyer of third-party data, the items that depend on what your supplier tells you are:

  • The sources or owners of the datasets, and how the datasets further the intended purpose of the system.
  • The number of data points, which may be given in general ranges, with estimates for dynamic datasets.
  • The types of data points, and whether the datasets include data protected by copyright, trademark or patent, or are entirely in the public domain.
  • Whether the datasets were purchased or licensed by the developer.
  • Whether the datasets include personal information or aggregate consumer information as defined in the CCPA (Civil Code § 1798.140).
  • Whether there was cleaning, processing or other modification, and its intended purpose.
  • The time period during which the data was collected, including a notice if collection is ongoing, and the dates the datasets were first used in development.
  • Whether the system used synthetic data generation.

The statute exempts systems used solely for security and integrity, for operating aircraft in national airspace, and national-security or defense systems made available only to a federal entity [1]. It has no standalone penalty clause [1]; commentators expect enforcement through California's Unfair Competition Law, so treat an inaccurate summary as a consumer-protection exposure. The practical failure mode is a fine-tuning release where the team cannot say whether a licensed corpus contained personal information because the license file never recorded it. The fine-tuning provider and developer duties page covers when a fine-tune makes you the developer.

Colorado SB26-189: documentation that flows to deployers

Colorado's training-data duty is a business-to-business documentation duty, not a public posting duty. SB26-189 replaced the consumer protections of SB 24-205 when it was signed on May 14, 2026 [3]. From January 1, 2027, a developer of ADMT that materially influences a consequential decision must give deployers technical documentation covering intended uses, categories of training data and known limitations, among other items [3].

Two practical points follow. First, "categories of training data" is coarser than AB 2013's list, so one well-built supplier record can feed both. Second, the Colorado Attorney General released interim draft rules on October 6, 2026 and is taking written comments through October 26, 2026 [4]. The documentation format may tighten when rules are final, so treat any template you build now as provisional.

Colorado's privacy statute applies separately when consequential-decision data includes personal data; see the Colorado Privacy Act and AI data licensing page rather than folding it into this row.

Laws that govern AI use, not training data

Most state AI activity in 2025 and 2026 regulates decisions, interactions or safety, and should sit in a separate tracker column. A multi-state roundup from the period groups California's employment regulations on automated decision systems with an Illinois law requiring disclosure of AI use in employment decisions [6]. Neither asks what data trained the model; both ask how the model's output is used and disclosed to workers or applicants.

California's SB 53 is the closest edge case. It applies to frontier models trained with more than 10^26 computational operations and requires safety frameworks and transparency reports [5]. Those reports may reference data-related risk controls, but SB 53 is not a training-data disclosure statute, and counsel should not cite it as one.

Sector rules also matter. Insurance regulators address external consumer data used in AI systems through the NAIC model bulletin and state adoptions; see insurance regulators on external data for AI. Privacy statutes reach training data whenever it contains personal information; the CCPA risk-assessment page and CCPA and licensing business data cover that ground. Recorded calls bring in two-party consent rules covered under federal and state wiretap laws.

No federal training-data disclosure statute

As of October 2026, no federal statute requires developers to disclose training data. The U.S. Copyright Office's Part 3 report on generative AI training is still a pre-publication version from May 2025 and is analysis, not a disclosure mandate [8]. Pending federal and state bills should be tracked in a separate "watch" list, never in the enacted table. For the wider US picture, see the United States law overview.

How to keep a state training-data tracker accurate

A tracker is only reliable if every row carries a status date, a primary-source link and a reviewer. The Colorado transition shows why: one public compliance tracker ranking for Colorado queries still listed SB 24-205 as the operative law in early June 2026, weeks after SB26-189 was signed [7][3]. Rows built from secondary summaries drift silently.

Use these rules:

  • Link the official bill page (leginfo.legislature.ca.gov, leg.colorado.gov) as the primary source, and the regulator page (for example, coag.gov) for rulemaking status.
  • Record "enacted", "effective" and "rules final" as separate dates; they rarely coincide.
  • Keep a column for "reaches training data: yes / indirect / no" so adjacent laws do not inflate scope.
  • Re-check each row on a fixed cadence and on every model release or substantial modification.
  • Store the supplier evidence each row depends on next to the row, not in a separate legal folder.

Illustrative example: invented to show structure; it does not describe an available dataset.

tracker_row:
  state: CO
  law: SB26-189
  status_checked: 2026-10-09
  enacted: 2026-05-14
  core_effective: 2027-01-01
  rules_status: interim draft 2026-10-06; comments to 2026-10-26
  reaches_training_data: yes   # developer-to-deployer documentation
  record_required: [intended_uses, training_data_categories, known_limitations]
  primary_source: https://leg.colorado.gov/bills/sb26-189
supplier_record_fields_feeding_this_row:
  dataset_id: DS-0142
  source_owner: "US logistics company (supplier of record)"
  data_point_types: [support_tickets, agent_notes]
  record_count_range: "100k-1M"
  collection_period: 2021-03 to 2025-12, not ongoing
  personal_information: removed_or_replaced; method logged; sample checked
  copyright_status: owned by supplier; licensed to buyer
  acquisition: licensed
  processing: deduplication, PII replacement, language filter
  synthetic_data_used: false
  first_used_in_development: 2026-02-10
  known_limitations: "English only; two product lines"

The same supplier record answers AB 2013's list and Colorado's categories, and it supports your retention schedule for training data and its records.

What to ask data suppliers so state disclosures are possible

The fastest way to fail a state disclosure is to buy data without the fields the statute names. Before signing, ask suppliers for: the owner of each source, record counts or ranges, collection dates, whether personal or aggregate consumer information is present and how it was treated, intellectual-property status, and what processing was applied. Put those answers in the license schedule or a provenance file delivered with the data; see the data provenance buyer's guide.

When you source through SourceX, each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which counsel can draw on when preparing AB 2013 and Colorado records. Buyers can describe the data they need at SourceX for buyers.

Getting training data with records for state AI laws

SourceX sources operational datasets from US companies on request and manages licensing, with each dataset rights-reviewed and diligence materials prepared for it. Nothing is contracted until a supplier agrees, and a request does not guarantee a match. Start a request at SourceX for buyers.

More from the AI training data compliance hub and the EU AI Act Article 53 obligations page.

Frequently asked questions

Does AB 2013 apply if we only fine-tune an open-weight model?

It can. AB 2013 covers substantial modifications, and commentators read that to include fine-tuning and retraining [2]. If you make the fine-tuned system available to Californians, document the data you added.

Is Colorado SB 24-205 still relevant?

Only historically. SB26-189 replaced its consumer protections in May 2026 [3]. Trackers and vendor guides that still cite SB 24-205 duties should be corrected [7].

Do state privacy laws count as training-data laws?

Not on their own, but they apply whenever training data contains personal information. Track them in a privacy column and link the relevant analysis, such as the CCPA risk-assessment rules, rather than duplicating them.

Sources

  1. California Legislature (California Legislative Information), "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  2. Conventus Law, "US: California's AB 2013 Requires Generative AI Data Disclosure By January 1, 2026". https://conventuslaw.com/report/us-californias-ab-2013-requires-generative-ai-data-disclosure-by-january-1-2026/
  3. Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
  4. Colorado Attorney General, "Colorado Automated Decision-Making Technology & Chatbot Safety Rulemaking" (2026). https://coag.gov/ai/
  5. Morrison & Foerster LLP, "At the Frontier - California Enacts AI Safety and Transparency Regulation TFAIA (SB 53)" (2025). https://www.mofo.com/pdf/resources/insights/251001-california-enacts-ai-safety-transparency-regulation-tfaia-sb-53
  6. Seyfarth Shaw LLP, "Artificial Intelligence Legal Roundup: Colorado Postpones Implementation of AI Law as California Finalizes New Employment Discrimination Regulations and Illinois Disclosure Law Set to Take Effect". https://www.seyfarth.com/news-insights/artificial-intelligence-legal-roundup-colorado-postpones-implementation-of-ai-law-as-california-finalizes-new-employment-discrimination-regulations-and-illinois-disclosure-law-set-to-take-effect.html
  7. Layer3 Labs, "AI Law Compliance Tracker" (2026). https://www.layer3labs.io/guides/ai-law-compliance-tracker
  8. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data