Regulation and governance for data buyers
Training Data Risk Assessment: Scoring a Dataset Before You Acquire It
Quick answer
A training data risk assessment scores a candidate dataset before purchase on eight dimensions: rights clarity, personal-data presence, sector rules, high-risk use, general-purpose AI disclosure exposure, export and national-security transfer rules, cross-border transfer, and supplier reliability. Each dimension gets a 0–3 score with evidence attached. The highest single score, not the average, sets the tier, and the tier decides whether counsel, privacy or security must sign off. Record the result in your data register and re-run it whenever the license or intended use changes.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why an intake score belongs before signature, not after delivery
The assessment has to happen before signature because the cheapest fix for most dataset risk is a contract term, a scope cut or a walk-away, and none of those exist after delivery. HITRUST's draft AI security requirements make the point directly: an organization should identify and evaluate the legal, regulatory and contractual obligations attached to its AI data [1]. NIST's AI RMF 1.0 frames the same work under its MAP function (context and risk identification) and GOVERN function (accountability and policy) [2], and the generative AI profile, NIST AI 600-1, lists data privacy and intellectual property among the twelve risks generative systems create or amplify [3].
A score also gives you a defensible record. When an auditor, a customer's security reviewer or a regulator asks why a dataset was approved, you can point to the evidence and the reviewer who signed off rather than reconstruct a decision from emails. Keep the scoring separate from quality and acceptance testing, which belong in steps like an evaluation of a fine-tuning dataset before buying.
The eight dimensions and what evidence moves each score
Each dimension should be scored only from documents the supplier provides, not from its sales description. If a supplier cannot produce evidence, score the dimension as if the risk were present. Use the rubric below as a starting template and adapt the thresholds to your policy.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Dimension | 0 (low) | 1 | 2 | 3 (high) | Evidence to request |
|---|---|---|---|---|---|
| Rights clarity | Supplier created the data; employee or contractor IP assignment on file | Supplier owns most content; minor third-party material identified | Mixed origin; some content licensed in with unclear sublicensing | Origin unknown, scraped, or acquisition method undocumented | Chain-of-title memo, contractor agreements, upstream licenses |
| Personal-data presence | No personal data by design | Pseudonymized with documented method | Direct identifiers removed but free text unreviewed | Raw names, emails, account numbers or biometrics | De-identification method, sample check results, field inventory |
| Sector rules | No regulated sector | Regulated sector, non-covered records | Covered records with documented de-identification (HIPAA, FERPA) | Covered records, no de-identification basis; children's data | Sector determination, Safe Harbor or Expert Determination report |
| High-risk use | Internal tooling | Customer-facing assistant | Model feeds consequential decisions in a review loop | Annex III use case or consequential decision without human review | Use-case description, model card draft |
| GPAI disclosure exposure | Not used in a general-purpose or public generative model | Fine-tuning only, small share of training content | Material share of a public generative model's training data | Primary corpus for a GPAI model placed on the EU market | Source description fields for AB 2013 and the EU summary template |
| Export and national-security transfer | No sensitive categories; no transfer to countries of concern | Sensitive categories below bulk thresholds | Bulk sensitive data, US-only access | Bulk sensitive data with any access path for a country of concern | Data category map, access list, vendor and employee locations |
| Cross-border transfer | Data stays in one jurisdiction | Transfer under an adequacy decision | Transfer under SCCs with a transfer impact assessment | No transfer mechanism identified | Transfer mapping, SCC module, TIA |
| Supplier reliability | Established supplier, audited controls, prior clean delivery | Established, unaudited | New supplier, partial documentation | Reseller that cannot name the original source | DDQ answers, security review, references |
Rights clarity carries the most weight in most portfolios because copyright and contract claims attach to the data itself and follow it into every model trained on it. For acquisition-method risk specifically, see lawful access and pirated sources; for code corpora, add a copyleft contamination check under the same dimension.
Regulatory triggers that raise a dimension automatically
Some facts should move a dimension to 2 or 3 regardless of how clean the rest of the file looks. These triggers are status as of October 2026 and should be re-checked at each review.
- EU high-risk use. Article 10 of the AI Act requires training, validation and testing data for high-risk systems to be subject to data governance and management practices covering collection processes, data origin, preparation and bias examination [4]. Regulation (EU) 2026/1744, published on 24 July 2026, amended the AI Act, including Article 10 [5]; high-risk application dates moved to 2 December 2027 for Annex III and 2 August 2028 for Annex I. See EU AI Act Article 10 governance requirements.
- General-purpose AI providers. Article 53 obliges GPAI providers to maintain a copyright policy that honors text and data mining reservations under Article 4(3) of Directive (EU) 2019/790 and to publish a summary of training content [6]. The Code of Practice copyright chapter describes how signatories document that policy [7]. A dataset headed for a GPAI model that cannot supply source descriptions for the summary scores 3 on disclosure exposure.
- California AB 2013. Developers of generative AI made available to Californians had to post training data documentation by January 1, 2026, including sources, whether data includes personal information and whether it was purchased or licensed [8]. Any dataset feeding such a system needs those fields captured at intake.
- Colorado. The Colorado Attorney General states SB26-189 takes effect January 1, 2027, with interim draft rules released October 6, 2026 and still open for comment [9]. Score automated consequential-decision uses affecting Colorado residents at 2 or higher on high-risk use until the final rules are settled.
- Children's data. The amended COPPA Rule had a compliance date of April 22, 2026 [10]. Any record set that may contain data collected from children under 13 scores 3 on sector rules until the collection basis is documented.
- Bulk sensitive personal data. The DOJ's rule at 28 CFR Part 202 restricts or prohibits data transactions that give countries of concern access to bulk US sensitive personal data, including genomic, biometric, health, financial and precise geolocation categories [11]. Map every vendor, annotator and cloud region with access before you score this dimension.
- EU personal data in models. EDPB Opinion 28/2024 remains the reference on whether a model trained on personal data is anonymous and how unlawfully processed training data can affect later use of the model [12]. The GDPR changes proposed in the Digital Omnibus are not law as of October 2026, so do not score against them.
Turning scores into tiers and sign-offs
The tier should follow the maximum dimension score, because one unresolved 3 can sink a deal that otherwise averages 1. Averaging hides exactly the single fatal defect the assessment exists to catch. Use the table below to route reviews.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Tier | Rule | Required sign-off | Typical outcome |
|---|---|---|---|
| Tier 1 | All dimensions 0–1 | Governance lead | Approve; standard license terms |
| Tier 2 | Any dimension at 2 | Governance lead plus the owner of that dimension (counsel for rights, privacy for personal data, security for transfer) | Approve with conditions written into the license or a remediation step |
| Tier 3 | Any dimension at 3, or three or more at 2 | Counsel, privacy and security jointly; executive risk owner | Remediate and re-score, narrow the scope, or decline |
Conditions at Tier 2 should be concrete: a removal of identified fields, a written representation on upstream licenses, a restriction to evaluation use, or a geographic access limit. Supplier-level tiering is a separate exercise; the data supplier risk tiering guide covers which vendors need a full review, and the due diligence questionnaire supplies the questions that generate evidence for this rubric.
Recording the result in your data register
Every assessment should produce one register entry that a later reviewer can read without the original team. Store the scores with the evidence references, the reviewer names, the conditions and the re-assessment triggers. The record also becomes the source for AB 2013 documentation and the EU training-content summary, so capture those fields once.
Illustrative example: invented to show structure; it does not describe an available dataset.
assessment_id: DRA-2026-0412
dataset: "B2B support ticket histories, 2019-2025, English"
intended_use: [fine_tuning, eval]
model_scope: "internal support assistant; not a GPAI model"
scores:
rights_clarity: 1 # supplier-authored tickets; customer attachments excluded
personal_data: 2 # names and emails replaced; free-text review sample pending
sector_rules: 0
high_risk_use: 1
gpai_disclosure: 0
export_natsec_transfer: 1 # no bulk sensitive categories identified
cross_border: 2 # EU-hosted annotator; SCCs plus TIA
supplier_reliability: 1
tier: 2
sign_offs: {governance: "J. Rivera", privacy: "A. Chen", counsel: null}
conditions:
- "Complete free-text identifier review on a 2,000-ticket sample before delivery"
- "License restricts use to named models; no redistribution"
evidence: [chain_of_title.pdf, deid_method_v2.pdf, tia_2026-09.pdf, ddq_response.xlsx]
reassess_on: [license_amendment, new_use_case, new_model_scope, regulatory_change]
next_review: 2027-04-01
When to re-assess an approved dataset
Re-assess whenever the license, the use or the law changes, because each one can move a dimension without anything happening to the data. A license amendment can narrow permitted uses; a decision to move from fine-tuning a narrow assistant to pretraining a general-purpose model changes GPAI disclosure exposure from 0 to 3. Regulatory dates also move, as the 2026 amendment to the AI Act showed [5].
Set a calendar review at least annually and tie the event triggers to your change process so a new use case cannot reach training without a fresh score. The governance policy template shows where those triggers sit in a written policy, and the NIST AI RMF guide for third-party data maps the cycle to MANAGE. For the evidence auditors expect to see afterward, use audit readiness for training data.
Where scoring meets sourcing
A risk assessment is only as good as the evidence the supplier can produce, so source from parties that document rights and preparation. SourceX finds US companies that hold operational datasets you describe, such as support and sales histories, engineering records and finance or legal workflows, and manages the commercial process from licensing to ongoing purchases. Every dataset is reviewed for ownership and consents, personal details like names, emails and account numbers are removed or replaced with the method recorded and a sample checked, and diligence materials covering source, rights, preparation and allowed use are prepared per dataset. No de-identification method is perfect, so your own privacy score still applies. Buyers can describe the data they need on the SourceX buyers page; data is sourced on request, and a request does not guarantee a match.
For broader context, start at the AI training data compliance hub, the full AI training data due diligence checklist and the guide on how to procure enterprise training data.
Get datasets that arrive with the evidence your assessment needs
SourceX assesses each candidate dataset's data and licensing permissions before anything is agreed, and every dataset is delivered under a license that defines records, uses, term and delivery. Every release is approved by the supplying company, and nothing is contracted until a supplier agrees. Tell SourceX what data your team needs.
Sources
- HITRUST (via Manula), "Identify and evaluate AI compliance and legal obligations (HITRUST AI security certification requirements, draft)". https://www.manula.com/manuals/hitrust/ai-security-certification-requirements-draft/1/en/topic/id-and-evaluate-ai-compliance-legal-obligations
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)" (2024). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- Official Journal of the European Union (EUR-Lex), "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- Colorado Attorney General, "Colorado Automated Decision-Making Technology & Chatbot Safety Rulemaking" (2026). https://coag.gov/ai/
- Federal Trade Commission, Federal Register, "Children's Online Privacy Protection Rule (Final Rule amendments), 90 FR 16918" (2025). https://www.federalregister.gov/documents/2025/04/22/2025-05904/childrens-online-privacy-protection-rule
- Federal Register / Department of Justice, "Prohibiting Transactions Involving Certain Bulk Sensitive Personal Data and Government-Related Data (90 FR 1636)" (2025). https://www.federalregister.gov/documents/2025/01/08/2024-29789/prohibiting-transactions-involving-certain-bulk-sensitive-personal-data-and-government-related-data
- European Data Protection Board, "Opinion 28/2024 on the processing of personal data in the context of AI models" (2024). https://edpb.europa.eu/our-work-tools/our-documents/opinion-board-art-64/opinion-282024-processing-personal-data-context-ai_en
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.