Skip to content

Regulation and governance for data buyers

The GPAI Model Documentation Form: Completing the Training Data Fields

Quick answer

The Model Documentation Form in the GPAI Code of Practice's Transparency chapter asks providers to describe training, testing and validation data by modality, provenance, how it was obtained and selected, size, scope and main characteristics, curation methods, and the measures used to detect unsuitable sources and identifiable biases. Each answer is marked for the AI Office, national authorities or downstream providers. For licensed data, most answers come from supplier records you should request at contract time, not reconstruct later.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What the form is and where its data fields come from

The Model Documentation Form is the Transparency chapter's template for recording the information that Article 53(1)(a) and (b) of the AI Act require, and its data fields track Annex XI and Annex XII. The Commission published the Code on 10 July 2025 as a voluntary tool with three chapters: Transparency, Copyright, and Safety and Security [1][2]. Signing it is a way to demonstrate compliance, not a substitute for the Act itself [1].

Annex XI Section 1, point 2(c), which Article 53(1)(a) points to, is the source of almost every data field. It asks, where applicable, for the type and provenance of data and curation methodologies such as cleaning and filtering, the number of data points, their scope and main characteristics, how the data was obtained and selected, and all other measures to detect the unsuitability of data sources and methods to detect identifiable biases [3]. Annex XII sets out the narrower package owed to downstream providers that integrate the model into their own AI systems [3].

As of October 2026, Article 53 duties have applied since 2 August 2025, and the AI Office's enforcement powers apply from 2 August 2026 for new models [3]. Free and open-source models without systemic risk are exempt from the documentation duties in points (a) and (b), but not from the copyright policy or the public summary [3]. Check the form version you use against the current DOCX on the Commission's Code page, because field wording, not this page, controls [1].

The training data fields, one by one

The data section breaks Annex XI point 2(c) into discrete fields, and each one needs a specific answer type, not narrative. Treat the list below as a reading of the Annex mapped to the form's structure; confirm field labels against your copy.

  • Data type / modality. Text, images, audio, video, code, structured records, or other. List every modality that reached training, including modalities used only in post-training.
  • Provenance. Where each component came from: web crawling, publicly available datasets, private non-publicly available datasets obtained from third parties, user data, synthetic data, or other sources [4]. Licensed enterprise data belongs in the third-party private category.
  • How the data was obtained and selected. Collection method (crawler, API, licensed transfer, recording program), selection criteria, and inclusion rules [3].
  • Number of data points. Size per modality with the unit stated: tokens, documents, images, audio hours. The unit matters more than precision; mixed units are a common review finding.
  • Scope and main characteristics. Domains, languages, time span, geography, and known gaps [3].
  • Curation methodologies. Deduplication, quality filters, toxicity filters, PII removal, decontamination against evaluation sets [3].
  • Measures to detect unsuitable data sources. How you screened for unlawful, low-quality or out-of-policy sources, including rights reservations under the copyright policy [3].
  • Measures to detect identifiable biases. Audits performed, attributes examined, and what was done about findings [3].

The form also has model-level fields on training process and compute that sit outside the data section. Those are owned by the training team, not the data owner.

Which recipient sees which answer

The form tiers each answer by recipient, so the confidentiality decision is made field by field rather than for the document as a whole. Commentary describes three audiences: the AI Office, national competent authorities, and downstream providers, with general information shared downstream and more sensitive detail reserved for authorities [4][5].

A useful working rule follows from comparing the two Annexes. Annex XII's downstream package concerns what integrators need to meet their own obligations, so type, provenance and curation summaries are the likely downstream-facing items, while exact data point counts, selection mechanics and source-screening detail are Annex XI material held for authorities on request [3]. Where the form's own marking differs from this rule, follow the form.

That split shapes your supplier contracts. A licensor may be comfortable with its data being described as "licensed customer support transcripts, US, 2019 to 2024" in a downstream package, but not with its company name or record counts leaving the regulator tier. Agree on the descriptor wording before the dataset ships; renegotiating confidentiality after the model launches is slow.

Mapping each field to the supplier record that answers it

Every data field should trace to a named record you hold, and for licensed data most of those records originate with the supplier. Build the mapping during procurement so the form owner is filling fields from files, not from interviews with people who have since left.

Illustrative example: invented to show structure; it does not describe an available dataset.

Form fieldRecord that answers itWho produces itLikely tier
Data type / modalityDataset manifest (file types, schema, MIME counts)Supplier, verified on intakeDownstream + authorities
ProvenanceLicense agreement recital and data origin statementSupplier and buyer counselDownstream (category) + authorities (detail)
How obtained and selectedDelivery record plus buyer selection logSupplier (collection), buyer (selection)Authorities
Number of data pointsIntake count report in stated units (records, tokens after tokenization)Buyer data engineeringAuthorities
Scope and characteristicsDatasheet or dataset card: domains, languages, date rangeSupplier draft, buyer editDownstream + authorities
Curation methodologiesSupplier preparation log (de-identification method, redaction) plus buyer pipeline configBothDownstream summary + authorities detail
Unsuitable source detectionRights review memo, consent basis, opt-out handlingSupplier and buyer counselAuthorities
Bias detection measuresBuyer audit report on the training mixBuyerAuthorities

Two failure modes recur. First, counts reported by the supplier in records do not match counts in tokens after your tokenizer, deduplication and filtering; record both and state which one the form uses. Second, the supplier's de-identification method is described in the contract but not in a file, so the curation field ends up vague. A machine-readable dataset card, for example using the Croissant-RAI vocabulary for provenance and collection metadata, keeps these facts attached to the data itself [7]. For the card structure, see dataset cards for licensed enterprise data.

How the form differs from the public training content summary

The Model Documentation Form is not public: it is kept per model version, provided to the AI Office and national authorities on request, and its downstream tier is made available to integrators; the public summary under Article 53(1)(d) is published and follows a separate Commission template dated 24 July 2025 [5][6]. The public template sets a common minimal baseline for what everyone can see, while the form holds the detail that authorities and integrators need [6].

The practical consequence is consistency. Every category and source type in the public summary should reconcile to the form's provenance field, and every licensed dataset named or described in one should appear in the other. Our guide to completing the public summary for licensed and private datasets covers that template, and the comparison of AB 2013, the EU summary and Colorado disclosures shows how one supplier record can feed all three.

The form also differs from Annex IV technical documentation for high-risk systems. If a downstream integrator builds a high-risk system on your model, your Annex XII package feeds their Annex IV documentation of training data sets, which is why downstream-tier answers need to be specific enough to use.

Versioning and retention

Draw up the form for each model version and keep it for ten years after the model is placed on the market, according to commentary on the Transparency chapter [4]. A new fine-tune, continued pre-training run or changed data mix means a new version of the data section, even if the architecture fields do not change.

Retention creates a contract question for licensed data. Your license may require deletion of the dataset at term end, while the form must still describe it for a decade. Keep the descriptive records (manifest, counts, preparation log, rights memo) separate from the data, and confirm the license lets you retain them. The retention requirements guide covers that conflict, and regulator access to licensed datasets covers what the AI Office can ask to see.

A field-completion checklist for licensed data

Before you mark the data section complete, each licensed component should pass the checks below.

Illustrative example: invented to show structure; it does not describe an available dataset.

  1. Modality and provenance category recorded per dataset, not per training run.
  2. License reference (agreement ID, effective date, permitted uses) linked from the provenance field.
  3. Counts recorded in supplier units and in post-pipeline tokens, with the form unit stated.
  4. Supplier preparation method (for example, direct identifiers replaced with tokens, sample reviewed) captured as a file, not a sentence.
  5. Rights review memo and any consent basis on file, cross-referenced to the copyright policy under the Code's copyright chapter.
  6. Downstream-tier descriptor wording approved by the supplier.
  7. Public summary categories reconciled to the form.
  8. Records stored outside the dataset deletion path.

For the statutory background, see Article 53 training data obligations, the audit readiness evidence pack, and the compliance hub. If you are still negotiating, the pre-training license rights guide lists the grant terms that make these records available.

Where SourceX fits in documenting licensed training data

SourceX sources operational datasets from US companies on request, and every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which gives the form owner supplier-side records for the provenance, curation and source-suitability fields. Describe the data you need on the SourceX buyer page; a request does not guarantee a match.

Request licensed data with documentation

SourceX looks for US businesses that hold the data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is transacted. Personal details are removed or replaced before delivery, and the method is recorded. Start a request at sourcex.si/buyers.

Sources

  1. European Commission (AI Office), "The General-Purpose AI Code of Practice" (2025). https://digital-strategy.ec.europa.eu/en/policies/gpai-code-practice
  2. Freshfields, "The final General-Purpose AI Code of Practice: a short guide" (2025). https://technologyquotient.freshfields.com/post/102ksv0/the-final-general-purpose-ai-code-of-practice-a-short-guide
  3. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  4. Latham & Watkins, "EU AI Act: GPAI Model Obligations in Force and Final GPAI Code of Practice in Place" (2025). https://www.lw.com/en/insights/eu-ai-act-gpai-model-obligations-in-force-and-final-gpai-code-of-practice-in-place
  5. Bird & Bird, "Taking the EU AI Act to Practice: Decoding the GPAI Code of Practice and the Training Data Summary Template" (2025). https://www.twobirds.com/en/insights/2025/taking-the-eu-ai-act-to-practice-decoding-the-gpai-code-of-practice-and-the-training-data-summary-te
  6. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  7. Jain et al. (MLCommons), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data