Regulation and governance for data buyers
EU AI Act Article 53: The Training Data Obligations of General-Purpose AI Model Providers
Quick answer
Article 53 gives every provider of a general-purpose AI (GPAI) model four duties that reach training data: confidential technical documentation describing the data (Annex XI), information for downstream providers (Annex XII), a copyright policy that honors Article 4(3) DSM opt-outs, and a public summary of training content on the AI Office template [1][2][3]. The duties have applied since 2 August 2025 [8]. Licensed data must arrive with enough provenance to feed all four.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
The four Article 53 duties that touch training data
Article 53(1) lists four obligations, and each one needs a different slice of your data records [1]. Points (a) and (b) are documentation duties, while points (c) and (d) concern copyright and public transparency. The table maps each duty to its audience and the training-data facts it consumes.
| Article 53(1) point | What it requires | Audience | Training-data facts it needs |
|---|---|---|---|
| (a) Technical documentation | Draw up and keep up to date documentation including the training and testing process, with at least the Annex XI items | AI Office and national competent authorities, on request | Type and provenance, curation, size, characteristics, how obtained and selected, unsuitability and bias checks |
| (b) Downstream information | Give providers integrating the model information enabling them to understand capabilities and limits, with at least the Annex XII items | Downstream AI system providers | Training data type, provenance and curation, at a less granular level |
| (c) Copyright policy | Put in place a policy to comply with Union copyright law, including identifying and complying with Article 4(3) reservations through state-of-the-art technologies | Internal, evidenced to the AI Office | License chain, opt-out checks, lawful-access evidence |
| (d) Public summary | Make publicly available a sufficiently detailed summary of training content using the AI Office template | The public | Data source categories, licensed vs. scraped vs. user vs. synthetic, processing measures |
Providers established outside the EU must also appoint an authorised representative under Article 54 before placing the model on the EU market [2]. For that scenario, see placing a US-trained model on the EU market.
Annex XI: what the confidential documentation must say about data
Annex XI asks for a data description detailed enough that a regulator can see where the corpus came from and how it was shaped [2]. As of October 2026, the training-data item in Annex XI Section 1 covers, where applicable, the data used for training, testing and validation, including the type and provenance of data and curation methodologies (cleaning, filtering), the number of data points, their scope and main characteristics, how the data was obtained and selected, and measures to detect unsuitable data sources and identifiable biases. Verify the exact wording against the consolidated text on EUR-Lex before you build templates on it.
The failure mode is a corpus description that stops at "licensed proprietary data." For each licensed dataset, your internal record should hold the supplier's record count, collection window, language and domain mix, file formats (for example JSONL conversation turns, PDF or HTML documents, Parquet tables), the de-identification method applied, and every filtering step your pipeline ran after receipt. If you cannot reconstruct "how obtained and selected" from a contract and a manifest, the Annex XI entry will be thin.
Annex XI documentation is supplied to the AI Office and national competent authorities on request, and Article 53(7) treats information obtained under the article as confidential [1]. That confidentiality is what makes it workable to describe a licensed dataset in detail. Your license still has to permit the disclosure; a non-disclosure clause that forbids naming the dataset to "any third party" can conflict with a regulator request. The page on regulator access to licensed training datasets covers that clause in depth.
Annex XII: what downstream providers receive
Annex XII is the lighter, shareable layer: downstream providers get enough about training data to understand the model's limits and meet their own duties [2]. As of October 2026, it includes, where applicable, information on the data used for training, testing and validation, including type, provenance and curation methodologies. It does not require the counts, selection logic or bias measures that Annex XI asks for.
In practice, Annex XII material often ships as a model card or integration guide given to enterprise customers under NDA. Because it leaves your control, licensed-data descriptions in it should match the categorized wording you use in the public summary, not the dataset-level detail of Annex XI. If a downstream provider fine-tunes your model and becomes a provider itself, the analysis on fine-tuning provider duties under the EU AI Act and AB 2013 applies.
Article 53(1)(c): the copyright policy and Article 4(3) opt-outs
The copyright policy must identify and comply with reservations of rights made under Article 4(3) of Directive (EU) 2019/790, including through state-of-the-art technologies [1]. Article 4(3) lets rightsholders opt out of the general text and data mining exception in Article 4, the one commercial developers rely on, so any web-derived content in a corpus needs an opt-out check at the time of collection. Licensed data moves the question from "was there a machine-readable reservation?" to "did the licensor have the right to grant TDM and training use?"
The GPAI Code of Practice Copyright chapter gives signatories a way to show compliance with Article 53(1)(c) [5]. Its measures include drawing up, keeping up to date and implementing a single copyright policy covering all GPAI models placed on the EU market, reproducing and extracting only lawfully accessible content, honoring machine-readable reservations such as robots.txt when crawling, mitigating infringing outputs, and designating a point of contact for rightsholders [5][7]. The Code is voluntary and adherence does not by itself establish compliance with copyright law [4].
For licensed corpora, the evidence that feeds the policy is contractual. Keep the executed license, the licensor's ownership and consent representations, the permitted-use clause naming training, and a record of any third-party content inside the dataset (embedded email attachments, quoted vendor documents, scanned contracts). The copyright chapter page on licensed training data and the guide to EU TDM opt-outs under DSM Article 4 go further.
Article 53(1)(d): the public training-content summary
The public summary must be "sufficiently detailed" and must follow the template the AI Office published on 24 July 2025 [3]. The Commission describes the template as a common minimal baseline for public disclosure, intended to help parties with legitimate interests, including copyright holders, exercise their rights [3][6]. Results report that the template groups content by source category, including commercially licensed content, web-scraped content, user-generated data and synthetic data, followed by data processing measures such as opt-out handling [6]. Check field-level requirements in the template file itself, since search summaries do not reproduce them in full.
Licensed datasets usually land in the "private, non-publicly available" category. Confirm with the licensor what the summary may say: the modality, the broad domain (for example, B2B customer-support transcripts or engineering tickets), the time range, and whether the data was licensed from rightsholders. The SourceX guide on what buyers need from suppliers for the EU training data summary lists supplier inputs field by field.
Open-source models and the carve-out
Free and open-source GPAI models get a partial exemption: points (a) and (b) do not apply when the model is released under a free and open-source license with publicly available weights, architecture and usage information, unless it is a model with systemic risk [1]. Points (c) and (d) still apply. Results report that the public summary duty covers open-source models too [6].
An open-weights release therefore still needs the copyright policy and the template summary. Teams that skip Annex XI recordkeeping because they publish weights often find later that they cannot fill the template's processing section either. Keep the dataset-level records regardless of release model.
Deadlines and enforcement as of October 2026
GPAI obligations under Article 53 have applied since 2 August 2025, and the AI Office's enforcement powers began on 2 August 2026 for new models [8]. Models placed on the market before 2 August 2025 have until 2 August 2027 to comply, under the transition in Article 111(3) [2][6]. Models with systemic risk carry additional Article 55 duties on top of Article 53.
Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026, amended the AI Act [9]. As of October 2026, commentary reports that it moved high-risk application dates, and the GPAI duties in Article 53 continue to apply; check the consolidated text listed on EUR-Lex for current wording [2][9]. For the high-risk side, see EU AI Act Article 10 data governance and deadlines after the Digital Omnibus.
License terms that let you meet all four duties
A training-data license for a GPAI provider should expressly allow the disclosures Article 53 requires, because a supplier NDA written for ordinary commercial deals can block them. The checklist below is a practical starting point for counsel and procurement. If you are sourcing operational data from businesses, you can also describe the dataset you need to SourceX, which rights-reviews each dataset and delivers it under a license defining records, uses, term and delivery.
Illustrative example: invented to show structure; it does not describe an available dataset.
Article 53 license and supplier record checklist
- Disclosure to authorities (Annex XI): license permits describing the dataset, its supplier and its processing to the AI Office and national competent authorities on request.
- Downstream description (Annex XII): license permits a categorized description (type, provenance, curation) in model documentation shared with integrators.
- Public summary (53(1)(d)): licensor approves the exact category wording for the template, for example "licensed B2B support transcripts, US, 2019-2024, English."
- Training use grant (53(1)(c)): permitted-use clause names model training, including pre-training and post-training, and states whether derived model weights are restricted.
- Ownership and consents: licensor representations on ownership, employee and customer consents, and third-party content embedded in records.
- Provenance manifest: record count, collection window, source systems (for example Zendesk exports, Jira issues, SharePoint documents), formats and schema version.
- Preparation record: de-identification method, sample check results, and filters applied before delivery.
- Retention: license lets you keep the manifest, license and preparation record for as long as the model is on the market and documentation must be kept current, even after data deletion obligations apply.
Illustrative example: invented to show structure; it does not describe an available dataset.
dataset_record:
internal_id: lic-2026-017
source_category_for_summary: "private, licensed from rightsholder"
modality: text
domain: "B2B SaaS customer support tickets and agent replies"
collection_window: "2020-01 to 2025-12"
records: "<supplier-stated count>"
formats: [jsonl]
obtained_via: "commercial license, executed agreement"
selection: "English tickets with resolved status; spam and auto-replies removed"
deidentification: "names, emails, phone and account numbers replaced; sample checked"
opt_out_check: "not applicable (non-web, licensed source)"
annex_xi_fields_covered: [type, provenance, curation, size, characteristics, obtained_selected, bias_checks]
The same record can feed the AB 2013 posting and other disclosure regimes; the disclosure requirements comparison shows the overlap. For the full regulatory map, start at the AI training data compliance hub or the SourceX overview of the EU AI Act for AI data licensing.
Sourcing licensed data for Article 53 documentation
SourceX sources operational datasets from US companies on request and manages the licensing process, serving AI teams wherever they are based. Every dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery, and diligence materials on source, rights, preparation and allowed use are prepared per dataset. Describe the data your model needs.
Frequently asked questions
Does Article 53 require naming every licensor publicly?
No requirement to name each licensor appears in Article 53(1)(d) itself; the duty is a sufficiently detailed summary on the AI Office template [1][3]. How granular the licensed-data section must be is set by the template and its explanatory notice, so check that file and agree the wording with each licensor.
Is signing the GPAI Code of Practice mandatory?
No. The Code is a voluntary tool, and providers that do not sign it must show compliance with Article 53 by other adequate means [4][7].
Does fine-tuning an existing GPAI model trigger Article 53?
It can, depending on the extent of the modification. The analysis is covered on the fine-tuning provider duties page.
Sources
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Parliament and Council of the European Union (EUR-Lex), "Regulation (EU) 2024/1689 (Artificial Intelligence Act)" (2024). https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- European Commission (AI Office), "The General-Purpose AI Code of Practice" (2025). https://digital-strategy.ec.europa.eu/en/policies/gpai-code-practice
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- WilmerHale, "European Commission Releases Mandatory Template for Public Disclosure of AI Training Data" (2025). https://wilmerhale.com/en/insights/blogs/wilmerhale-privacy-and-cybersecurity-law/european-commission-releases-mandatory-template-for-public-disclosure-of-ai-training-data
- Freshfields, "The final General-Purpose AI Code of Practice: a short guide" (2025). https://technologyquotient.freshfields.com/post/102ksv0/the-final-general-purpose-ai-code-of-practice-a-short-guide
- Latham & Watkins, "EU AI Act: GPAI Model Obligations in Force and Final GPAI Code of Practice in Place" (2025). https://www.lw.com/en/insights/2025/09/eu-ai-act-gpai-model-obligations-in-force-and-final-gpai-code-of-practice-in-place
- European Parliament and Council of the European Union (EUR-Lex), "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.