Skip to content

Regulation and governance for data buyers

The GPAI Code of Practice Copyright Chapter: Writing a Copyright Policy That Covers Licensed Training Data

Quick answer

The copyright chapter of the EU General-Purpose AI Code of Practice gives signatories five measures for meeting AI Act Article 53(1)(c): a single written copyright policy, lawful-access and anti-piracy rules for crawling, compliance with machine-readable rights reservations, mitigation of infringing outputs, and a contact point with a complaints process [2][4]. Most measures target web crawling. Licensed, non-crawled data is governed by your licenses, so your policy must define how rights scope is evidenced, tracked and enforced [3][4].

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

The copyright chapter is a voluntary compliance route for Article 53(1)(c), not a statement of what EU copyright law permits. The Commission published the final Code on 10 July 2025 with three chapters: Transparency, Copyright, and Safety and Security [1]. Article 53(1)(c) itself requires GPAI providers to put in place a policy to comply with Union copyright law and, in particular, to identify and comply with rights reservations under Article 4(3) of the DSM Directive (EU) 2019/790, including through state-of-the-art technologies [5].

Signing the Code is a way to demonstrate the policy duty; it does not by itself establish that any particular training use is lawful [1][2]. Non-signatories still owe the Article 53(1)(c) policy and must show compliance by other means. For the full set of GPAI data duties, see our guide to Article 53 training data obligations and the EU AI Act owner page.

The five measures, read against licensed data

Three of the five measures apply to every source; the two crawling measures reach licensed data only indirectly, or when part of the licensed corpus was itself crawled. Law-firm analyses of the final text count five measures, one fewer than the prior draft [4][7]. The table below maps each measure to its practical scope for a provider that mixes crawled web data with licensed corpora.

Measure (final chapter)Core commitmentCrawled web dataLicensed, non-crawled data
Copyright policyDraw up, keep up to date and implement one written policy for all GPAI models placed on the EU market [2][4]Governs crawl rulesMust contain the licensed-data rules (this is where diligence now lives)
Lawful accessReproduce and extract only lawfully accessible content; do not circumvent technical protection measures such as paywalls or subscription walls; exclude websites recognized as persistently infringing [4][7]Primary targetIndirect: confirms the licensor had lawful access and that the corpus is not pirated
Rights reservationsIdentify and honor machine-readable Article 4(3) reservations, including robots.txt and other appropriate protocols [4][5]Primary targetGenerally replaced by the license grant, unless the licensor's own content was assembled by crawling
Infringing outputsTechnical safeguards against reproducing protected content in outputs, plus a prohibition of infringing uses in terms of use [4]AppliesApplies equally; licensed works can be regurgitated too
Contact point and complaintsDesignate a point of contact for rightsholders and an electronic complaints mechanism [4]AppliesApplies; complaints about licensed works need a fast path to the license file

Two drafting consequences follow. First, a rightsholder complaint does not distinguish crawled from licensed sources, so your complaints workflow needs to resolve any work to its acquisition channel. Second, the output-mitigation measure is source-agnostic: a license to train rarely includes a right to reproduce the licensed text verbatim in model outputs.

Why non-crawled data is now your policy's problem

The final Code dropped a draft commitment to obtain adequate information about non-crawled training content, so the standard of diligence for licensed and private data is whatever your own policy says it is [4]. That is freedom, but it is also exposure. A complaint, a regulator query, or litigation discovery will test the policy you wrote against the records you actually hold.

The chapter also states that it does not affect agreements between signatories and rightsholders authorizing the use of works [3]. In practice the license is the operative rights instrument for licensed data, and the copyright policy should say how licenses are reviewed, recorded and enforced. The commitments are proportionate to provider size, taking account of SMEs and start-ups [3], which matters for a smaller fine-tuning team that inherits provider duties (see fine-tuning and provider duties).

Acquisition channel is also a live litigation variable. In July 2026 a court gave final approval to a class settlement in Bartz v. Anthropic [8]. A settlement is not a merits ruling, but it shows why a policy should record lawful access for every bulk corpus, licensed or not. Our page on lawful access and pirated sources covers that risk in depth.

The policy must be one document covering all GPAI models you place on the EU market; publishing a summary is encouraged but not required [4]. Counsel usually draft it as a controlled policy with an owner, a version history and annexes for operational detail. The outline below is a working structure, not text from the Code.

Illustrative example: invented to show structure; it does not describe an available dataset.

SectionWhat it must sayEvidence kept
1. Scope and modelsWhich GPAI models and versions the policy covers; placing-on-market datesModel register entries
2. Acquisition channelsCrawling, licensed datasets, user or customer data, public-domain and openly licensed corpora, synthetic dataChannel taxonomy mapped to the training-content summary categories
3. Licensed-data rulesRequired license fields, approval workflow, prohibited sources, expiry and termination handlingLicense register, rights record per dataset (see below)
4. Crawling rulesLawful access, no circumvention of paywalls or protection measures, excluded infringing domains, robots.txt and other reservation protocolsCrawler config snapshots, opt-out detection logs
5. Output mitigationRegurgitation filters, memorization testing, acceptable-use termsTest reports, terms-of-use version
6. Contact point and complaintsElectronic channel, triage, response steps, link from a work to its source recordComplaint log with outcomes
7. RolesPolicy owner, data acquisition lead, legal reviewer, crawl engineering ownerRACI
8. Review cadenceTriggers: new model, new channel, new protocol, material complaintReview minutes
9. Version historyChange log retained as audit evidenceSigned versions

Keep section 2 aligned with the public training-content summary required by Article 53(1)(d), which uses the AI Office template published 24 July 2025 [5][6]. If the summary says licensed private datasets were used, the policy should explain how those datasets were rights-reviewed. See completing the training content summary for licensed and private datasets and the GPAI model documentation form.

Drafting the licensed-data section

The licensed-data section should require a rights record for every licensed dataset before it enters a training pipeline. Licensed data does not need robots.txt checks at ingestion, but it does need evidence of rights scope on file: permitted training use, territory, term, field of use and what happens at expiry. Without that record the complaints measure cannot be operated, because nobody can answer "under what authority was this work used?"

Specify these controls in the section:

  • Rights scope fields. Grant of training use (pre-training, post-training, evaluation, retrieval), territory, term, field of use, sublicensing, and whether outputs may reproduce licensed text.
  • Upstream provenance. A licensor statement on how its content was assembled, and whether any portion was itself crawled; if so, that portion falls back under the crawling rules.
  • Personal data and third-party rights. Records of consent basis, redaction or de-identification method, and excluded fields; copyright compliance does not settle GDPR questions.
  • Expiry and termination. What happens to checkpoints trained before expiry, and how the dataset is quarantined from future training runs.
  • Pipeline binding. A dataset ID that travels with every shard into training manifests, so a complaint can be traced to a license.

Illustrative example: invented to show structure; it does not describe an available dataset.

dataset_id: LIC-2026-0417
acquisition_channel: licensed_non_crawled
content: support ticket threads and resolution notes, English, 2019-2025
licensor_role: originating business (holds the records)
license_ref: LA-0417 v2 (executed)
rights_scope:
  permitted_uses: [pre_training, post_training, evaluation]
  retrieval_or_display: not_permitted
  territory: worldwide
  term_end: 2029-06-30
  output_reproduction: not_permitted
upstream_provenance:
  crawled_portion: none (licensor attestation on file)
  third_party_works_inside: customer attachments excluded at export
personal_data:
  method: names, emails, phones, account numbers replaced with typed tokens
  sample_check: recorded
on_expiry: exclude from new runs; legal review for existing checkpoints
complaint_route: copyright-contact queue -> license register lookup
policy_section: 3.2

Crawling rules and rights reservations, briefly

For crawled data, the policy should name the reservation protocols you honor and how you detect them. Article 53(1)(c) requires identifying and complying with Article 4(3) reservations using state-of-the-art technologies [5], and the Code points signatories to robots.txt and other appropriate machine-readable protocols [4][7]. Record crawler user-agent strings, the date each robots.txt was fetched, and how reservations found later are applied to data already collected.

Failure modes to address in writing: honoring robots.txt only at first crawl and not on re-crawls; treating a missing robots.txt as consent when the site's terms or a TDMRep header reserve rights; and mixing third-party crawl snapshots whose filtering you cannot verify. Our comparison of robots.txt, ai.txt, TDMRep and other AI usage signals covers detection tooling.

Keeping the policy consistent with other records

The copyright policy, the public summary and the technical documentation must tell the same story about sources. Regulators and rightsholders will compare them. A practical check is quarterly reconciliation: every acquisition channel in the summary appears in policy section 2, every licensed dataset in the summary's private-data disclosure has a rights record, and every complaint closed in the period cites a source record.

For providers training outside the EU, the policy still applies to models placed on the EU market; see EU copyright duties for models trained outside the EU. For a broader internal governance document that the copyright policy can sit beside, use the AI training data governance policy template, and browse the rest of the compliance hub.

Where a licensed sourcing process supports the policy

A licensed sourcing process supports section 3 of the policy when each dataset arrives with a license and a diligence record. SourceX sources operational datasets from US companies on request, rights-reviews each one for ownership and consents, and delivers it under a license defining records, uses, term and delivery. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which maps directly to the rights record above. SourceX does not source scraped web content, so these datasets sit in your licensed, non-crawled channel. Buyers can describe the data they need on the buyer request page; a request does not guarantee a match. See also our source policy.

SourceX sources operational datasets from US companies on request, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, with diligence materials prepared per dataset. Describe the data you need at sourcex.si/buyers.

Frequently asked questions

Does the copyright chapter require robots.txt checks for licensed datasets?

No measure requires opt-out checks on content you license directly from a rightsholder; the rights reservation measure concerns crawling [4][5]. If the licensor assembled part of its corpus by crawling, treat that part under your crawling rules.

Must the copyright policy be published?

The policy must exist as one written, maintained and implemented document; publishing a summary of it is encouraged but not required [4]. The training-content summary under Article 53(1)(d) is a separate, mandatory publication [5][6].

Does a license override the Code's commitments?

The chapter states it does not affect agreements between signatories and rightsholders authorizing use of works [3]. The license governs the scope of use, while the policy still governs output mitigation and complaints for that content.

Sources

  1. European Commission (AI Office), "The General-Purpose AI Code of Practice" (2025). https://digital-strategy.ec.europa.eu/en/policies/gpai-code-practice
  2. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
  3. European Commission (copy hosted by Edition Multimedia), "Code of Practice for General-Purpose AI Models: Copyright Chapter (10 July 2025)" (2025). https://www.editionmultimedia.fr/wp-content/uploads/2025/07/Code-of-Practice-for-GPAI-Copyright-10-07-25.pdf
  4. Slaughter and May (The Lens), "Mining the copyright chapter of the GPAI Code" (2025). https://thelens.slaughterandmay.com/post/102ktcs/mining-the-copyright-chapter-of-the-gpai-code
  5. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  6. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  7. A&O Shearman, "EU Artificial Intelligence Office publishes the final version of the GPAI Code of Practice" (2025). https://www.aoshearman.com/en/insights/ao-shearman-on-data/eu-artificial-intelligence-office-publishes-the-final-version-of-the-gpai-code-of-practice
  8. Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data