Skip to content

Regulation and governance for data buyers

The EU Training Content Summary Template: Completing It for Licensed and Private Datasets

Quick answer

For commercially licensed and privately obtained datasets, the EU public summary of training content asks for a categorised, narrative description rather than a list of every contract or work. Under the AI Office template published 24 July 2025, you describe licensed data by modality, content type, origin and scale in the private-datasets part of the data-sources section, keep it consistent with your copyright policy, and use the explanatory notice to decide when a licensor or dataset must be named and when a general description is enough [1][3][4].

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What the template asks for and where licensed data sits

The template has three parts: general information about the model, a list of data sources by type and origin, and data processing aspects [3]. Article 53(1)(d) of the AI Act requires providers of general-purpose AI (GPAI) models to publish a "sufficiently detailed summary" of training content using the AI Office template, and the Commission describes the template as a common minimal baseline for what goes public [1][2]. As of October 2026, the Article 53 duties have applied since 2 August 2025, and the AI Office's enforcement powers apply from 2 August 2026 for new models [2].

Licensed and private data belongs in the data-sources section, which separates publicly available datasets, private datasets obtained from third parties, crawled and scraped data, user data, synthetic data and other sources [3][8]. Private third-party data is split again into data commercially licensed from rightsholders or their representatives and private data obtained from other third parties. Check the subsection numbering and wording against the current DOC template before you fill it in, because commentary paraphrases it in different ways [1].

For where this summary sits next to other duties, see the Article 53 obligations for GPAI providers and the side-by-side comparison of AB 2013, the EU summary and Colorado disclosures.

How much detail "sufficiently detailed" means for licensed data

"Sufficiently detailed" means comprehensive in scope, not technically exhaustive: the summary should let a rightsholder or data subject work out whether their kind of content was likely used, without listing every file [4]. Commentators note that the template aims to balance transparency with the protection of trade secrets and confidential business information [4]. In practice, the template asks for more granularity on scraped and public data than on licensed data, where rightsholders already know from the contract that their works may be used [1][6].

That asymmetry shows in the crawled-data part. There, providers describe crawler behavior and the collection period, and list the top 10% of scraped domain names by content size, with relaxations for SMEs (top 5% or 1,000 domains, whichever is lower) [6]. Licensed and private datasets are usually described in aggregate by modality, content category and source type. Confirm in the notice whether any threshold makes an individual private dataset nameable, for example when it is already publicly known [1].

Under-disclosure is the larger enforcement risk. The template is meant to help rightsholders and data subjects exercise their rights under copyright and data protection law, so a vague entry such as "proprietary data from partners" invites challenges [5]. Over-disclosure is a contract risk. Naming a supplier, record counts or a deal scope can breach a licence's confidentiality clause.

Naming licensors: what to disclose and what to keep out

Whether you must name individual licensors depends on the explanatory notice and your contracts, so treat naming as a decision to check rather than a default. The notice and template frame licensed data as a category to describe, and the summary is not an inventory of agreements [1][4]. Many providers therefore state that they licensed data from rightsholders, then describe the types of rightsholder (for example, news publishers or enterprise software companies), the modalities and the content domains.

Three checks settle most naming questions:

  • Is the dataset or licensor already public? If a licence has been announced, naming it adds little confidentiality risk and may be expected. Confirm the notice's wording on publicly known private datasets [1].
  • Does the licence permit categorised disclosure? Many data licences restrict naming the counterparty. Clause drafting belongs in your licensing playbook; this page only covers what the summary needs.
  • Does the description stay accurate in aggregate? "Licensed customer support transcripts from US software companies, English, 2019 to 2024" is useful and names no one.

Keep the description consistent with your copyright policy under Article 53(1)(c) and, if you have signed it, the Copyright chapter of the GPAI Code of Practice (published 10 July 2025) [2][7]. A summary that says data was "licensed" while the copyright policy relies on the text and data mining exception for the same corpus creates an inconsistency that rightsholders can use.

A worked entry for licensed and private datasets

A good licensed-data entry gives modality, content type, origin category, time span, language and scale band, plus a pointer to processing steps. The record below shows the internal fields a documentation owner can keep per dataset and the public text each row produces. It is designed so the public wording can be generated from the private record without exposing contract terms.

Illustrative example: invented to show structure; it does not describe an available dataset.

Internal field (kept private)Example valuePublic summary wording
source_categoryprivate_third_party / commercially_licensed"Data commercially licensed from rightsholders"
licensor_typeUS mid-market SaaS companies (3 licences)"Licensed from software companies"
licensor_namesWithheld under licence confidentialityNot disclosed; category only
modalitytext"Text"
content_typesupport tickets, agent replies, internal knowledge-base articles"Customer support conversations and help-center documentation"
languageen-US"Primarily English"
collection_window2018-01 to 2024-06"Content created between 2018 and 2024"
personal_data_handlingnames, emails, phone and account numbers replaced; method logged; sample checkedDescribed under data processing aspects
licence_scope_refcontract ID, permitted uses, termNot disclosed
used_inpre-training v2; fine-tuning v2.1Reflected in model version scope

A matching public paragraph might read: "We trained on text commercially licensed from software companies, consisting of customer support conversations and help-center documentation created between 2018 and 2024, primarily in English. Personal data was removed or replaced before training." That text satisfies the category-level expectation and does not disclose licensor names, contract terms or record counts.

For "private datasets obtained from other third parties", such as data from a broker or an affiliate, add the type of third party and how you acquired it. Do not stop at "licensed". Acquisition route matters for copyright exposure, as covered in lawful access and pirated sources.

Data processing aspects for private data

The data processing section asks how you respected rights reservations and handled illegal content and personal data, and private data needs entries there too [3]. For licensed corpora, state that rights were obtained by licence and that the copyright policy governs any opt-out or reservation handling. For personal data, describe the de-identification measures in general terms, such as removal or pseudonymization of direct identifiers, and avoid claiming anonymity unless your assessment supports it under GDPR Recital 26 and EDPB Opinion 28/2024.

Common failure modes in this section:

  • Copy-paste from the Annex XI technical documentation. The public summary and the GPAI model documentation form serve different audiences, so do not leak confidential fields into the public text.
  • Claims the supplier record cannot back. If a supplier cannot document consent basis or de-identification, do not describe the data as "fully anonymized".
  • Silence on fine-tuning data. Licensed instruction or preference data used in post-training is still training content.

Keeping the summary current across model versions

Update the summary when further training, fine-tuning or a new model version adds training content, and keep a version log that ties each summary to a model release [1]. Check the explanatory notice for the update trigger and cadence it sets. Also check the transition rules for models placed on the market before 2 August 2025 have until 2 August 2027 to publish their summaries [1][6].

A practical control is to make the internal per-dataset record (the table above) the single source for the public summary, the Annex IV technical documentation where it applies, and US disclosures such as California AB 2013. When a new licensed dataset arrives, its record should be complete before training starts. The inputs to request from suppliers are listed in what buyers need from suppliers for EU training data summaries. Broader context is on the EU AI Act and data licensing page and in the compliance hub for data buyers.

If your summary depends on licensed operational data, sourcing it with documentation in place makes the entries easier to write. SourceX for AI data buyers prepares diligence materials (source, rights, preparation and allowed use) for each dataset it sources.

Sourcing licensed data you can describe in your summary

SourceX sources operational datasets from US companies on request, rights-reviews each one for ownership and consents, and delivers it under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, and the method is recorded. Describe the data you need at SourceX for AI data buyers.

Frequently asked questions

Can I name my data supplier in the training content summary?

You can if your licence permits it, and you may be expected to if the dataset or deal is already public. Otherwise, a category-level description by rightsholder type, modality and content usually fits the template's approach. Confirm against the explanatory notice and the confidentiality clause [1][4].

Do synthetic data generated from licensed data and the licensed source both need entries?

Yes. Synthetic data has its own category in the data-sources section, and the licensed seed data remains training content if it was used directly [3][8]. Describe the generating model category and the seed data category.

Does the summary replace my copyright policy?

No. Article 53(1)(c) requires a separate copyright policy, and the summary should be consistent with it [2]. Rightsholders read both documents together [5].

Sources

  1. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  2. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  3. WilmerHale, "European Commission Releases Mandatory Template for Public Disclosure of AI Training Data" (2025). https://wilmerhale.com/en/insights/blogs/wilmerhale-privacy-and-cybersecurity-law/european-commission-releases-mandatory-template-for-public-disclosure-of-ai-training-data
  4. Shoosmiths, "EU Releases Template for AI Training Data Transparency: What It Means for Providers of General-Purpose AI Models" (2025). https://www.shoosmiths.com/insights/articles/ai-confessions-what-your-model-did-last-summer
  5. Digital Policy Alert, "European Commission's template for the public summary of training content for general-purpose artificial intelligence models" (2025). https://digitalpolicyalert.org/change/15553-european-commissions-template-for-the-public-summary-of-training-content-for-general-purpose-artificial-intelligence-models
  6. Slaughter and May, "Just in time: EU AI Office publishes template for summarising GPAI training content" (2025). https://thelens.slaughterandmay.com/post/102kyy4/just-in-time-eu-ai-office-publishes-template-for-summarising-gpai-training-conte
  7. European Commission (AI Office), "The General-Purpose AI Code of Practice" (2025). https://digital-strategy.ec.europa.eu/en/policies/gpai-code-practice
  8. datos.gob.es, "More transparency in AI: new template for documenting general-purpose model training data" (2025). https://datos.gob.es/en/blog/more-transparency-ai-new-template-documenting-general-purpose-model-training-data

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data