Regulation and governance for data buyers
Colorado SB 26-189: Training Data Documentation for Automated Decision-Making Technology
Quick answer
As of October 2026, Colorado SB 26-189 replaces the 2024 Colorado AI Act with a law built around automated decision-making technology (ADMT). From 1 January 2027, a developer whose ADMT materially influences a consequential decision must give deployers technical documentation covering intended uses, categories of training data, known limitations, and instructions for appropriate use and human review, and must tell deployers about material updates [1][2]. Vendors of hiring, lending, insurance and housing tools should build that training data record now, starting with what their data suppliers can document.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
How SB 26-189 replaced SB 24-205
SB 26-189 is the operative Colorado law for decision models, and SB 24-205's "high-risk AI system" framework no longer governs. The path was not direct: SB25B-004, passed in an August 2025 special session, pushed SB 24-205's start date from 1 February 2026 to 30 June 2026 [6]. Commentators then reported that enforcement was paused in April 2026 and that the act was replaced in May 2026 [7]. Governor Polis signed SB 26-189 on 14 May 2026, and it appears as Chapter 131 of the 2026 session laws [1][3].
The practical consequence is that many compliance checklists, vendor questionnaires and contract templates still cite SB 24-205 terms such as "high-risk AI system" and "algorithmic discrimination" duties [7]. If your deployer customers send questionnaires built on the old statute, answer against the new one and say so. For a wider view of state rules that reach training data, see US state AI laws that reach training data.
| Date | Event | What it means for training data records |
|---|---|---|
| Aug 2025 | SB25B-004 delays SB 24-205 to 30 Jun 2026 [6] | Old-framework documentation projects slowed |
| Apr 2026 | SB 24-205 enforcement reported paused [7] | Old questionnaires remain in circulation |
| 14 May 2026 | SB 26-189 signed, Chapter 131 [1][3] | ADMT framework replaces the old consumer protections |
| 6 Oct 2026 | AG releases interim draft rules; comments due 26 Oct 2026 [2] | Detail on documentation may change |
| 1 Jan 2027 | Core developer and deployer obligations start [1][2] | Documentation must exist and reach deployers |
The Division of Real Estate summary notes that some sections may carry different effective dates, so check the enrolled text section by section [4].
Which systems and decisions the law reaches
The law reaches ADMT that processes personal data to produce outputs, such as scores, rankings or predictions, that materially influence consequential decisions [8]. In practice that is the resume ranker, the credit risk score, the underwriting tier or the tenant screening score a deployer relies on when deciding about a person [1][8]. Whether you are a "developer" generally turns on building or substantially modifying the technology rather than hosting it; confirm against the enrolled definitions [1].
Two scoping questions decide how much documentation work you face. First, does the output "materially influence" the decision, or is it one weak input among many that a human overrides routinely; the answer belongs in your intended-use statement, not in a sales deck. Second, does the model process personal data at inference time even if it was trained on de-identified data; most consequential-decision tools do. The legislative fiscal note is a useful neutral reading of scope while the AG rulemaking runs [5].
What the deployer documentation must contain
From 1 January 2027, the developer gives each deployer technical documentation with four content blocks, plus notice of material updates [1]. The official summary lists:
- Intended uses: the decisions, populations and settings the system was designed and validated for, and uses you do not support.
- Categories of training data: what kinds of data the model learned from (see the next section).
- Known limitations: performance gaps, populations or conditions where outputs are less reliable, and known failure modes.
- Instructions for appropriate use and human review: how deployers should interpret outputs, when a human should review or override, and what the system should not be used to decide alone.
- Material updates: deployers must be notified when the system changes materially [1].
The AG's interim draft rules, released on 6 October 2026, may add format or content detail, and written comments are due by 26 October 2026 [2]. Treat any field list you build now as provisional until final rules publish. Commentators also report that developers and deployers must keep records proving compliance for at least three years [9]; align that with your license deletion duties using training data retention requirements.
What "categories of training data" should cover
The statute's phrase is "categories," which signals a description by type rather than a record-level inventory, but the exact level of detail must be checked against the enrolled text and the final AG rules [1][2]. A defensible reading is that a deployer should be able to tell, from your description, what populations, time periods, sources and outcome definitions shaped the model. "Proprietary and third-party data" fails that test; "2019–2024 loan application and repayment records from US consumer lenders, with charge-off at 24 months as the outcome label" passes it.
For decision models, five category dimensions carry most of the weight:
- Source type: first-party operational records, licensed third-party records, bureau or consortium data, synthetic data, or public data.
- Population and time window: who is represented, from where, and over which years, including known under-representation.
- Feature families: application fields, transaction history, employment history, claims history, free text, and any location or device fields that can act as proxies for protected characteristics.
- Label definition: what "good hire," "default" or "claim fraud" means, and whether the label records a past human decision or an observed outcome.
- Preparation: de-identification, sampling, reweighting, exclusions and imputation applied before training.
Label definition is where decision models most often go wrong. When the label is a historical approval, hire or claim decision, the model learns the earlier decision-makers' patterns, including their biases; see historical decision bias in operational labels. For lending, outcome labels tied to credit decision records with adverse action reasons let you document both the decision and its stated basis.
A training data category record you can maintain per source
A per-source record, filled at acquisition and updated on every retrain, is the simplest way to produce Colorado documentation without a scramble. The same record also feeds California AB 2013 postings for generative systems [10] and EU AI Act Article 10 data governance evidence for high-risk systems [11].
Illustrative example: invented to show structure; it does not describe an available dataset.
training_data_category:
record_id: TDC-0007
system: "tenant-screen-score v3.2"
category_name: "Residential lease applications and payment outcomes"
source_type: licensed_third_party # first_party | licensed_third_party | synthetic | public
supplier_record_ref: "license-2026-014, schedule A"
population: "US rental applicants, 9 states; under-represents rural applicants"
time_window: "2020-01 to 2025-06"
feature_families: [application_fields, rent_payment_history, income_verification_flags]
proxy_risk_fields: [zip5, prior_address_count] # reviewed for protected-class correlation
label:
name: "lease_default_12m"
type: observed_outcome # observed_outcome | past_human_decision
definition: "90+ days arrears or eviction filing within 12 months of move-in"
preparation:
deidentification: "names, SSNs, phone numbers and account numbers replaced with tokens"
exclusions: "applications without a move-in date"
reweighting: none
known_limitations: "no applicants who were declined, so outcome is unobserved for them"
allowed_uses_per_license: "model training and validation for tenant screening"
deployer_summary_text: "Lease application and 12-month payment outcome records, 2020-2025, nine US states"
last_reviewed: 2026-09-30
material_change_log: []
The deployer_summary_text field is what goes into the deployer document; the other fields are your evidence behind it. The known_limitations entry above shows a classic reject-inference gap that also belongs in the limitations block of the deployer documentation.
What to require from training data suppliers
Your Colorado documentation is only as good as what your suppliers can tell you, so the obligations flow upstream into procurement. Ask each supplier for the following before signing, and put the delivery of these items into the license schedule.
| Ask the supplier for | Why it matters under SB 26-189 | Common failure |
|---|---|---|
| Plain description of source systems and record types | Basis for "categories of training data" [1] | "Aggregated industry data" with no source type |
| Population, geography and date range | Supports limitations and intended-use scoping | Date range given for extraction, not for the events |
| Field dictionary with definitions | Lets you flag proxy fields and feature families | Coded columns with no codebook |
| Outcome and decision field definitions | Label definition drives bias analysis | Approval flag mislabeled as performance outcome |
| De-identification method and sample check | Personal data handling and limitation disclosure | Method undocumented, so it cannot be described |
| Rights and consent basis, allowed uses | Confirms the data may be used for this purpose | Licence silent on decision-model training |
| Change notice when refreshed data differs | Feeds your material-update notices [1] | Silent schema or population drift between drops |
Insurers face parallel scrutiny of external data; the NAIC bulletin on third-party data for AI covers that overlay. Colorado's separate consumer privacy statute may also apply to personal data you license, subject to its exemptions; see the Colorado Privacy Act and AI data licensing.
When a training data change counts as a material update
A change in training data is likely material when it would change what you tell a deployer about intended uses, categories, limitations or appropriate use [1]. Until the AG rules define the trigger, set internal thresholds and apply them consistently. Reasonable triggers include:
- adding or removing a training data category or supplier;
- changing a label definition or the outcome window;
- extending the population to a new geography, product line or applicant segment;
- retraining that measurably shifts performance for a subgroup you report on;
- discovering a limitation, such as a proxy field or a reject-inference gap, not previously disclosed.
Log each decision, including changes you judged not material, with the evidence behind it, and keep the log under the three-year record practice commentators describe [9].
How the Colorado record lines up with other disclosure regimes
One well-structured supplier record can serve several regimes, though each asks different questions. California AB 2013 requires public website documentation of training data for generative AI systems and services [10]; Colorado's document goes privately to deployers of decision systems [1]. EU AI Act Article 10 asks providers of high-risk systems to apply data governance and quality criteria to training, validation and test sets [11]. The comparison is worked through in AI training data disclosure requirements compared and EU AI Act Article 10 data governance.
Where SourceX fits for decision-model training data
SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages the licensing and ongoing purchases. Data is sourced on request rather than held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, delivered under a license defining records, uses, term and delivery, and comes with diligence materials on source, rights, preparation and allowed use. Those materials can feed your category record; you can describe the decision data you need by type, population and outcome.
Request documented training data for Colorado-covered decision models
Describe the operational records your hiring, lending, insurance or housing model needs, and SourceX looks for US businesses that hold them; every release is approved by the supplying company. Nothing is contracted until a supplier agrees, and personal details such as names and account numbers are removed or replaced before delivery. Start at SourceX for AI data buyers, or browse the compliance guides for data buyers and the AI data hub.
Sources
- Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
- Colorado Attorney General, "Colorado Automated Decision-Making Technology & Chatbot Safety Rulemaking" (2026). https://coag.gov/ai/
- Office of the Governor of Colorado, "Governor Polis Signs Bills into Law Making Colorado An Even Better Place to Do Business" (2026). https://governorsoffice.colorado.gov/governor/news/governor-polis-signs-bills-law-making-colorado-even-better-place-do-business-breaking-down
- Colorado Division of Real Estate, "SB26-189 Summary" (2026). https://dre.colorado.gov/sb26-189-summary
- Colorado Legislative Council Staff, "SB 26-189 Fiscal Note" (2026). https://leg.colorado.gov/bill_files/115880/download
- Greenberg Traurig LLP, "Colorado Delays Comprehensive AI Law With Further Changes Anticipated" (2025). https://www.gtlaw.com/en/insights/2025/9/colorado-delays-comprehensive-ai-law-with-further-changes-anticipated
- The Employer Report, "AI Regulation on Hold in Colorado, but Employer Risk Isn't" (2026). https://www.theemployerreport.com/2026/05/ai-regulation-on-hold-in-colorado-but-employer-risk-isnt/
- Monitaur, "Colorado's SB26-189: New Requirements for Automated Decision-Making Technology and Compliance Strategies" (2026). https://www.monitaur.ai/blog-posts/colorados-sb26-189-new-requirements-for-automated-decision-making-technology-and-compliance-strategies
- Promise Legal, "Colorado AI Act Compliance: SB 26-189 Developer & Deployer Guide" (2026). https://blog.promise.legal/colorado-ai-act-compliance-guide/
- California Legislative Information, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.