Skip to content

Regulation and governance for data buyers

Training Data Risk Assessment: Scoring a Dataset Before You Acquire It

Quick answer

A training data risk assessment scores a candidate dataset before purchase on eight dimensions: rights clarity, personal-data presence, sector rules, high-risk use, general-purpose AI disclosure exposure, export and national-security transfer rules, cross-border transfer, and supplier reliability. Each dimension gets a 0–3 score with evidence attached. The highest single score, not the average, sets the tier, and the tier decides whether counsel, privacy or security must sign off. Record the result in your data register and re-run it whenever the license or intended use changes.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why an intake score belongs before signature, not after delivery

The assessment has to happen before signature because the cheapest fix for most dataset risk is a contract term, a scope cut or a walk-away, and none of those exist after delivery. HITRUST's draft AI security requirements make the point directly: an organization should identify and evaluate the legal, regulatory and contractual obligations attached to its AI data [1]. NIST's AI RMF 1.0 frames the same work under its MAP function (context and risk identification) and GOVERN function (accountability and policy) [2], and the generative AI profile, NIST AI 600-1, lists data privacy and intellectual property among the twelve risks generative systems create or amplify [3].

A score also gives you a defensible record. When an auditor, a customer's security reviewer or a regulator asks why a dataset was approved, you can point to the evidence and the reviewer who signed off rather than reconstruct a decision from emails. Keep the scoring separate from quality and acceptance testing, which belong in steps like an evaluation of a fine-tuning dataset before buying.

The eight dimensions and what evidence moves each score

Each dimension should be scored only from documents the supplier provides, not from its sales description. If a supplier cannot produce evidence, score the dimension as if the risk were present. Use the rubric below as a starting template and adapt the thresholds to your policy.

Illustrative example: invented to show structure; it does not describe an available dataset.

Dimension0 (low)123 (high)Evidence to request
Rights claritySupplier created the data; employee or contractor IP assignment on fileSupplier owns most content; minor third-party material identifiedMixed origin; some content licensed in with unclear sublicensingOrigin unknown, scraped, or acquisition method undocumentedChain-of-title memo, contractor agreements, upstream licenses
Personal-data presenceNo personal data by designPseudonymized with documented methodDirect identifiers removed but free text unreviewedRaw names, emails, account numbers or biometricsDe-identification method, sample check results, field inventory
Sector rulesNo regulated sectorRegulated sector, non-covered recordsCovered records with documented de-identification (HIPAA, FERPA)Covered records, no de-identification basis; children's dataSector determination, Safe Harbor or Expert Determination report
High-risk useInternal toolingCustomer-facing assistantModel feeds consequential decisions in a review loopAnnex III use case or consequential decision without human reviewUse-case description, model card draft
GPAI disclosure exposureNot used in a general-purpose or public generative modelFine-tuning only, small share of training contentMaterial share of a public generative model's training dataPrimary corpus for a GPAI model placed on the EU marketSource description fields for AB 2013 and the EU summary template
Export and national-security transferNo sensitive categories; no transfer to countries of concernSensitive categories below bulk thresholdsBulk sensitive data, US-only accessBulk sensitive data with any access path for a country of concernData category map, access list, vendor and employee locations
Cross-border transferData stays in one jurisdictionTransfer under an adequacy decisionTransfer under SCCs with a transfer impact assessmentNo transfer mechanism identifiedTransfer mapping, SCC module, TIA
Supplier reliabilityEstablished supplier, audited controls, prior clean deliveryEstablished, unauditedNew supplier, partial documentationReseller that cannot name the original sourceDDQ answers, security review, references

Rights clarity carries the most weight in most portfolios because copyright and contract claims attach to the data itself and follow it into every model trained on it. For acquisition-method risk specifically, see lawful access and pirated sources; for code corpora, add a copyleft contamination check under the same dimension.

Regulatory triggers that raise a dimension automatically

Some facts should move a dimension to 2 or 3 regardless of how clean the rest of the file looks. These triggers are status as of October 2026 and should be re-checked at each review.

  • EU high-risk use. Article 10 of the AI Act requires training, validation and testing data for high-risk systems to be subject to data governance and management practices covering collection processes, data origin, preparation and bias examination [4]. Regulation (EU) 2026/1744, published on 24 July 2026, amended the AI Act, including Article 10 [5]; high-risk application dates moved to 2 December 2027 for Annex III and 2 August 2028 for Annex I. See EU AI Act Article 10 governance requirements.
  • General-purpose AI providers. Article 53 obliges GPAI providers to maintain a copyright policy that honors text and data mining reservations under Article 4(3) of Directive (EU) 2019/790 and to publish a summary of training content [6]. The Code of Practice copyright chapter describes how signatories document that policy [7]. A dataset headed for a GPAI model that cannot supply source descriptions for the summary scores 3 on disclosure exposure.
  • California AB 2013. Developers of generative AI made available to Californians had to post training data documentation by January 1, 2026, including sources, whether data includes personal information and whether it was purchased or licensed [8]. Any dataset feeding such a system needs those fields captured at intake.
  • Colorado. The Colorado Attorney General states SB26-189 takes effect January 1, 2027, with interim draft rules released October 6, 2026 and still open for comment [9]. Score automated consequential-decision uses affecting Colorado residents at 2 or higher on high-risk use until the final rules are settled.
  • Children's data. The amended COPPA Rule had a compliance date of April 22, 2026 [10]. Any record set that may contain data collected from children under 13 scores 3 on sector rules until the collection basis is documented.
  • Bulk sensitive personal data. The DOJ's rule at 28 CFR Part 202 restricts or prohibits data transactions that give countries of concern access to bulk US sensitive personal data, including genomic, biometric, health, financial and precise geolocation categories [11]. Map every vendor, annotator and cloud region with access before you score this dimension.
  • EU personal data in models. EDPB Opinion 28/2024 remains the reference on whether a model trained on personal data is anonymous and how unlawfully processed training data can affect later use of the model [12]. The GDPR changes proposed in the Digital Omnibus are not law as of October 2026, so do not score against them.

Turning scores into tiers and sign-offs

The tier should follow the maximum dimension score, because one unresolved 3 can sink a deal that otherwise averages 1. Averaging hides exactly the single fatal defect the assessment exists to catch. Use the table below to route reviews.

Illustrative example: invented to show structure; it does not describe an available dataset.

TierRuleRequired sign-offTypical outcome
Tier 1All dimensions 0–1Governance leadApprove; standard license terms
Tier 2Any dimension at 2Governance lead plus the owner of that dimension (counsel for rights, privacy for personal data, security for transfer)Approve with conditions written into the license or a remediation step
Tier 3Any dimension at 3, or three or more at 2Counsel, privacy and security jointly; executive risk ownerRemediate and re-score, narrow the scope, or decline

Conditions at Tier 2 should be concrete: a removal of identified fields, a written representation on upstream licenses, a restriction to evaluation use, or a geographic access limit. Supplier-level tiering is a separate exercise; the data supplier risk tiering guide covers which vendors need a full review, and the due diligence questionnaire supplies the questions that generate evidence for this rubric.

Recording the result in your data register

Every assessment should produce one register entry that a later reviewer can read without the original team. Store the scores with the evidence references, the reviewer names, the conditions and the re-assessment triggers. The record also becomes the source for AB 2013 documentation and the EU training-content summary, so capture those fields once.

Illustrative example: invented to show structure; it does not describe an available dataset.

assessment_id: DRA-2026-0412
dataset: "B2B support ticket histories, 2019-2025, English"
intended_use: [fine_tuning, eval]
model_scope: "internal support assistant; not a GPAI model"
scores:
  rights_clarity: 1          # supplier-authored tickets; customer attachments excluded
  personal_data: 2           # names and emails replaced; free-text review sample pending
  sector_rules: 0
  high_risk_use: 1
  gpai_disclosure: 0
  export_natsec_transfer: 1  # no bulk sensitive categories identified
  cross_border: 2            # EU-hosted annotator; SCCs plus TIA
  supplier_reliability: 1
tier: 2
sign_offs: {governance: "J. Rivera", privacy: "A. Chen", counsel: null}
conditions:
  - "Complete free-text identifier review on a 2,000-ticket sample before delivery"
  - "License restricts use to named models; no redistribution"
evidence: [chain_of_title.pdf, deid_method_v2.pdf, tia_2026-09.pdf, ddq_response.xlsx]
reassess_on: [license_amendment, new_use_case, new_model_scope, regulatory_change]
next_review: 2027-04-01

When to re-assess an approved dataset

Re-assess whenever the license, the use or the law changes, because each one can move a dimension without anything happening to the data. A license amendment can narrow permitted uses; a decision to move from fine-tuning a narrow assistant to pretraining a general-purpose model changes GPAI disclosure exposure from 0 to 3. Regulatory dates also move, as the 2026 amendment to the AI Act showed [5].

Set a calendar review at least annually and tie the event triggers to your change process so a new use case cannot reach training without a fresh score. The governance policy template shows where those triggers sit in a written policy, and the NIST AI RMF guide for third-party data maps the cycle to MANAGE. For the evidence auditors expect to see afterward, use audit readiness for training data.

Where scoring meets sourcing

A risk assessment is only as good as the evidence the supplier can produce, so source from parties that document rights and preparation. SourceX finds US companies that hold operational datasets you describe, such as support and sales histories, engineering records and finance or legal workflows, and manages the commercial process from licensing to ongoing purchases. Every dataset is reviewed for ownership and consents, personal details like names, emails and account numbers are removed or replaced with the method recorded and a sample checked, and diligence materials covering source, rights, preparation and allowed use are prepared per dataset. No de-identification method is perfect, so your own privacy score still applies. Buyers can describe the data they need on the SourceX buyers page; data is sourced on request, and a request does not guarantee a match.

For broader context, start at the AI training data compliance hub, the full AI training data due diligence checklist and the guide on how to procure enterprise training data.

Get datasets that arrive with the evidence your assessment needs

SourceX assesses each candidate dataset's data and licensing permissions before anything is agreed, and every dataset is delivered under a license that defines records, uses, term and delivery. Every release is approved by the supplying company, and nothing is contracted until a supplier agrees. Tell SourceX what data your team needs.

Sources

  1. HITRUST (via Manula), "Identify and evaluate AI compliance and legal obligations (HITRUST AI security certification requirements, draft)". https://www.manula.com/manuals/hitrust/ai-security-certification-requirements-draft/1/en/topic/id-and-evaluate-ai-compliance-legal-obligations
  2. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
  3. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)" (2024). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
  4. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  5. Official Journal of the European Union (EUR-Lex), "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
  6. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  7. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
  8. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  9. Colorado Attorney General, "Colorado Automated Decision-Making Technology & Chatbot Safety Rulemaking" (2026). https://coag.gov/ai/
  10. Federal Trade Commission, Federal Register, "Children's Online Privacy Protection Rule (Final Rule amendments), 90 FR 16918" (2025). https://www.federalregister.gov/documents/2025/04/22/2025-05904/childrens-online-privacy-protection-rule
  11. Federal Register / Department of Justice, "Prohibiting Transactions Involving Certain Bulk Sensitive Personal Data and Government-Related Data (90 FR 1636)" (2025). https://www.federalregister.gov/documents/2025/01/08/2024-29789/prohibiting-transactions-involving-certain-bulk-sensitive-personal-data-and-government-related-data
  12. European Data Protection Board, "Opinion 28/2024 on the processing of personal data in the context of AI models" (2024). https://edpb.europa.eu/our-work-tools/our-documents/opinion-board-art-64/opinion-282024-processing-personal-data-context-ai_en

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data