Regulation and governance for data buyers
Annex IV Technical Documentation: How to Describe Licensed Training Data Sets
Quick answer
Annex IV point 2(d) of the EU AI Act asks high-risk providers, where relevant, for datasheets describing training methodologies and the data sets used: a general description, provenance, scope and main characteristics, how the data was obtained and selected, labelling procedures and cleaning methods [1]. For licensed data, write one entry per data set that separates what the supplier delivered from what your team derived, and back each statement with a license, delivery manifest or pipeline log.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What Annex IV point 2(d) actually asks for
Point 2(d) requires a datasheet-style description of each training data set covering six things: general description, provenance, scope and main characteristics, how the data was obtained and selected, labelling procedures (supervised learning is the example) and data cleaning methods (outlier detection is the example) [1]. It sits inside point 2, the detailed description of the system's elements and development process, which Annex IV splits across nine sections in total [1]. The words "where relevant" are best read narrowly: for a system that trains models on data, plan to complete this section in full rather than rely on the qualifier.
Three neighboring points also touch licensed data. Point 2(a) asks for the methods and steps of development, including use of pre-trained systems or tools provided by third parties, so a licensed corpus used to fine-tune a vendor base model shows up in two places. Point 2(g) covers validation and testing procedures, including the validation and testing data used and their main characteristics [1]. Write 2(d) and 2(g) from the same data inventory so they cannot contradict each other.
Annex IV documents the result; Article 10 data governance sets the quality bar that result must meet [3]. In practice, Article 10(2) lists the governance practices (design choices, collection processes and origin, preparation operations, assumptions, availability and suitability, bias examination, gaps), and your 2(d) entry is where an assessor checks that each practice actually happened [2].
Timing as of October 2026
As of October 2026, the high-risk obligations behind Annex IV are not yet in application, but the evidence you need comes from acquisitions you are making now. Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026, amended the AI Act, including Article 10 [5]; secondary commentary reports high-risk dates of 2 December 2027 for Annex III systems and 2 August 2028 for Annex I products. Check the consolidated text on EUR-Lex before you fix internal deadlines [1], and see Digital Omnibus planning dates.
The practical consequence is simple. A data set licensed in 2026 without a provenance record, a collection description or a labelling protocol is hard to document in 2027, because the supplier relationship, the people who ran the labelling and the original export may no longer be reachable.
Supplier-delivered facts versus buyer-derived facts
Every sentence in a 2(d) entry should be attributable to either the supplier or your own pipeline, and the entry should say which. Market surveillance authorities and, where involved, notified bodies can test claims against evidence, and the evidence for "collected from a ticketing system between 2019 and 2024" lives in the supplier's diligence pack, while the evidence for "deduplicated with MinHash at a 0.85 Jaccard threshold" lives in your repo.
Split each entry into two blocks:
- As received. Source system and export format, collection period, record unit and counts at delivery, the license reference and permitted uses, the supplier's de-identification method, any labels the supplier created and their guidelines, and known gaps the supplier disclosed.
- As prepared. Your filters, deduplication, language identification, PII rescans, train/validation/test splits, augmentation or synthetic expansion, relabelling, and the resulting counts at each step.
This split is also what protects you in a dispute. If a supplier's representation about consents turns out to be wrong, a 2(d) entry that clearly says "per supplier attestation dated X" is defensible; one that states the fact as your own finding is not. The dataset delivery specification template is a good place to require the "as received" fields contractually, before delivery. If you are scoping a new acquisition, you can also describe these documentation needs up front through SourceX's buyer intake.
Mapping datasheets and dataset cards to Annex IV
Existing datasheets and dataset cards are good raw material for 2(d), but none of them maps one-to-one to the Annex IV headings. Gebru et al.'s Datasheets for Datasets organizes questions under motivation, composition, collection process, preprocessing/cleaning/labeling, uses, distribution and maintenance [7]; Google's Data Cards add modular, audience-specific summaries [8]. Commentary on the developer duties in Colorado's original SB 24-205 (since replaced by SB26-189) likewise treated model cards and dataset cards as acceptable documentation vehicles [9], which suggests one well-kept card can feed several regimes.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Annex IV 2(d) element [1] | Datasheet section [7] | Evidence to attach | Typical gap in licensed data |
|---|---|---|---|
| General description | Motivation, composition | Supplier data description, record schema | Marketing summary instead of schema-level description |
| Provenance | Collection process | License, chain-of-title or ownership attestation, source-system names | Reseller cannot name the originating business process |
| Scope and main characteristics | Composition | Field dictionary, date range, language and geography distribution, class balance | Counts given only for the full corpus, not per split |
| How obtained and selected | Collection process, preprocessing | Export query or filter logic, sampling method, your inclusion criteria | Supplier pre-filtered records without stating the rule |
| Labelling procedures | Preprocessing/labeling | Annotation guidelines version, annotator qualifications, agreement metric (e.g. Cohen's kappa) | Labels inherited from operational fields with no guideline |
| Cleaning methods | Preprocessing/cleaning | Dedup method and threshold, outlier rules, PII scrub method and sample QA result | Cleaning done in notebooks with no versioned record |
Keep the card in a versioned repository next to the data manifest, and pin the 2(d) text to a specific card version and data snapshot hash. For deeper background on the card format, see SourceX's guide to dataset cards for licensed enterprise data and the glossary entry for a datasheet for datasets.
A worked 2(d) entry for a licensed support corpus
A good 2(d) entry reads like a lab record: specific enough that a reviewer could request the underlying artifact for every claim. The example below shows the shape for a licensed customer-support ticket corpus used to train a classifier component of a system assumed, for the example, to fall under Annex III.
Illustrative example: invented to show structure; it does not describe an available dataset.
dataset_id: DS-TRAIN-017
role: training # training | validation | testing (cross-reference 2(g))
snapshot: sha256:4f1c…e9a # manifest hash of the exact files used
general_description: >
English-language support tickets (subject, body, agent replies, resolution code)
from a US B2B software vendor's ticketing system.
provenance:
origin: supplier operational system (ticketing export, JSONL)
license_ref: LIC-2026-031, schedule A (permitted uses: training and evaluation)
rights_evidence: supplier ownership and consent attestation, dated 2026-05-12
intermediary: data broker; originating business confirmed in diligence pack
scope:
period: 2020-01 to 2025-12
records_received: 412,300 tickets
geography: US customers; EU-resident records excluded by supplier filter F-2
obtained_and_selected:
supplier_side: tickets with status=closed and >=1 agent reply
buyer_side: removed auto-generated alerts (rule R-4), language-ID en>=0.9
labelling:
source: supplier resolution_code (operational, not annotated for ML)
buyer_relabel: 18,000-ticket subset, guideline v3.2, 2 annotators, kappa 0.81
cleaning:
pii: supplier replaced names, emails, phones, account numbers; buyer rescan
found residual identifiers in 0.4% of a 2,000-record sample, redacted
dedup: MinHash, Jaccard 0.85, 6.1% removed
outliers: tickets >20k tokens excluded
known_limitations: SMB customers over-represented; no voice-channel tickets
Notice what the entry avoids: unsupported adjectives such as "high quality" or "representative." Representativeness is an Article 10(3) criterion [3], and it belongs in your examination of the data against the intended purpose, documented with distributions, not asserted. See applying the representative, error-free and complete criteria to purchased data.
Validation and test sets need their own entries
Validation and test data are documented under point 2(g), and they deserve separate entries rather than a footnote to the training set [1]. Record how each split was drawn (random, temporal, by customer or site), the leakage controls between splits, and whether the test set came from the same supplier and period as training data. A test set carved from the same licensed export is a weaker independence claim than one from a different supplier or later period, and the entry should say so plainly.
If you bought a separate held-out set, record its license terms on evaluation use and how access is restricted to prevent contamination. Regulatory requirements for validation and test data covers sector rules that sit on top of Annex IV.
Special categories, health data and de-identification statements
Where a licensed set contains special-category or health data, the 2(d) cleaning section should name the de-identification standard actually applied, not a generic "anonymized." For US health records, that means stating whether the supplier used HIPAA Safe Harbor or Expert Determination under 45 CFR 164.514(b), and attaching the expert's report reference if the latter [10]. Neither method is the same as GDPR anonymization, so an EU assessor will still want your own re-identification analysis.
If you process special-category data to detect bias under Article 10(5), the conditions and safeguards are separate obligations [3]; document them in the Article 10(5) bias-detection record and cross-reference rather than duplicating.
Confidential documentation versus public summaries
Annex IV technical documentation is prepared for national competent authorities and notified bodies, not for publication, so supplier names, license terms and contract references can sit here. That is different from the public training-content summary that general-purpose model providers publish using the AI Office template under Article 53(1)(d) [6]. If you are both a GPAI provider and a high-risk provider, keep one internal inventory and generate two outputs: a detailed, confidential 2(d) entry and a coarser public summary.
Check your licenses for this. Many data agreements restrict disclosure of the supplier's identity or terms; you need a carve-out that permits disclosure to regulators and assessors. Regulator access to licensed training data explains what that clause should cover.
Retention and change control for dataset entries
Article 18 requires providers to keep the technical documentation at the disposal of national competent authorities for 10 years after the system is placed on the market or put into service [4]. That period usually outlasts the data license, so separate two things in your agreements: the right to keep and use the data, and the right to keep documentation, manifests and hashes describing it after the data is deleted.
Change control matters as much as retention. Every retraining on a refreshed supplier delivery creates a new snapshot and a new 2(d) entry version; tie these into the data management procedure in your quality management system, described in Article 17 data management procedures.
Pre-acquisition checklist for Annex IV-ready data
The cheapest time to secure Annex IV evidence is before signature. Ask each supplier for:
- A schema-level description and field dictionary for the export.
- The originating system and business process, plus ownership and consent attestations.
- The selection or filter logic applied before delivery.
- Labelling guidelines, if any labels are included, and who produced them.
- The de-identification method used, with sample QA results.
- Known gaps, biases and exclusions.
- License terms permitting disclosure to regulators and notified bodies, and retention of documentation beyond the data term.
For the wider regulatory map, start at the AI training data compliance hub.
Sourcing licensed data with documentation inputs
SourceX sources operational datasets from US companies on request and manages the licensing process; every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which gives your 2(d) entry concrete inputs, though a request does not guarantee a match. Describe the data you need at SourceX for AI data buyers.
Sources
- European Parliament and Council (EUR-Lex), "Regulation (EU) 2024/1689 (Artificial Intelligence Act)" (2024). https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
- Presencis, "EU AI Act Article 10: Data governance (conformity assessment)". https://presencis.com/regulations/eu-ai-act/conformity/article-10-data-governance
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- European Commission, AI Act Service Desk, "AI Act Article 18: Documentation keeping". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-18
- European Parliament and Council (EUR-Lex), "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- Gebru et al., arXiv, "Datasheets for Datasets" (2018). https://arxiv.org/abs/1803.09010v7
- Pushkarna, Zaldivar, Kjartansson, arXiv, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- IRMI, "Colorado Artificial Intelligence Law: Developer Disclosure Requirements". https://www.irmi.com/articles/expert-commentary/colorado-artificial-intelligence-law-developer-disclosure-requirements
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.