Skip to content

Regulation and governance for data buyers

Placing a US-Trained Model on the EU Market: Copyright Duties, Training Location and the Authorised Representative

Quick answer

Yes, the AI Act's copyright duties follow a general-purpose AI model into the EU market even when every training run happened in Virginia or Oregon. Article 53(1)(c) of the AI Act requires providers placing such models on the EU market to adopt a policy to comply with Union copyright law, including Article 4(3) text and data mining opt-outs [2]. Recital 106 says this applies regardless of where training took place [1]. Non-EU providers must also appoint an EU authorised representative first [2].

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What Recital 106 actually says about training location

Recital 106 states that providers placing general-purpose AI (GPAI) models on the Union market should comply with the copyright-policy obligation "regardless of the jurisdiction" in which the copyright-relevant acts underpinning training take place [1]. The stated rationale is a level playing field: no provider should gain a competitive advantage in the EU by training under lower copyright standards elsewhere [1]. Commentary reads it as a direct answer to forum shopping, where a model is trained in a permissive jurisdiction and then offered globally [1].

Two caveats matter for counsel. First, recitals are interpretive; they shape how the AI Office and courts read the articles, but they do not create freestanding obligations. Second, the binding hook is not the recital but the market-placement trigger in Article 53 [2]. The practical reading is that the EU regulates your policy and documentation as a condition of market access, not the act of copying in a US data center.

The operative obligation is a market-placement duty under Article 53(1)(c)

The duty that binds you is Article 53(1)(c): a GPAI provider must put in place a policy to comply with Union law on copyright and related rights, in particular to identify and comply with, including through state-of-the-art technologies, reservations of rights expressed under Article 4(3) of Directive (EU) 2019/790 [2]. The same article requires a sufficiently detailed public summary of training content under 53(1)(d), built on the Commission template published 24 July 2025 [2][5]. These obligations have applied since 2 August 2025; as of October 2026, the AI Office's enforcement powers apply to new models from 2 August 2026 [2].

Note what the text does not say. It does not require that training itself comply with EU copyright law as a matter of territorial infringement; it requires a policy that identifies and respects opt-outs. Academic commentary points out that this sits uneasily with the territorial nature of copyright [6]. The practical consequence is that a lab offering a model in the EU should honor EU opt-outs across its whole training pipeline, wherever compute runs [1].

Open-source release does not remove this obligation. Article 53 exempts qualifying free and open-source models from some documentation duties but keeps the copyright policy and training summary requirements in place [2]. A US lab releasing open weights that EU developers download should plan as if 53(1)(c) applies.

How the Code of Practice turns the policy into commitments

The GPAI Code of Practice Copyright chapter is the most concrete template for an Article 53(1)(c) policy. Signatories draw up, keep up to date and implement a single copyright policy covering all GPAI models they place on the EU market [3]. They commit to reproduce and extract only lawfully accessible content when crawling, to honor machine-readable rights reservations such as robots.txt, and to provide a point of contact and complaint mechanism for rightsholders [3][9].

The Code is voluntary, published 10 July 2025, and adherence is not by itself proof of compliance with Union copyright law [4]. For a non-EU provider, its main value is evidentiary: it gives the AI Office a recognized structure against which to read your policy. Our guide to writing a copyright policy that covers licensed training data maps each measure to records you can request from data suppliers.

Territoriality cuts both ways in litigation

Copyright infringement claims remain territorial even though the AI Act's policy duty is not. In UK litigation over a model trained on stock images, the primary training-infringement claims were withdrawn because there was no evidence that training took place in the UK [7]. That outcome shows that where compute ran is a factual question that can decide whether a national infringement claim survives.

The same logic means a US fair use defense does not answer EU questions. A June 2025 Northern District of California ruling found fair use on the record before it in a training case that remains ongoing [8]. Even if US training is ultimately excused under 17 U.S.C. § 107, a model placed on the EU market still needs a policy addressing Article 4(3) opt-outs [2]. For the US side, see what lawful access means for copyright risk.

Two failure modes recur in diligence:

  • Location claims without logs. A lab asserts "all training in the US" but cannot produce cloud region records, job IDs or cluster locations for every pre-training and fine-tuning run, including contractor runs.
  • Opt-out blind spots in licensed data. A licensed corpus is assumed clean, but the supplier's upstream collection included crawled content where no one checked TDM reservations.

Licensed data: align territory and field of use with every market

Licensed operational data is the cleanest answer to Article 4(3), because the rightsholder has expressly authorized the use instead of leaving you to detect an opt-out. The license grant still has to match where the model goes. A license limited to "use in the United States" or silent on territory creates an argument that EU market placement of a model trained on the data falls outside the grant.

Check four clauses before an EU launch:

  1. Territory: the grant covers training wherever compute runs and deployment of resulting models in the EU, or worldwide.
  2. Field of use: pre-training, fine-tuning, evaluation and commercial deployment are named, not implied.
  3. Rights warranty scope: the supplier confirms ownership and consents for the records delivered, and how any third-party content inside them was handled.
  4. Documentation cooperation: the supplier will provide the facts you need for the training content summary and for regulator requests [5].

Our related guides cover completing the EU training content summary for licensed datasets and the broader Article 53 training data obligations. For whether a US supplier can license to your lab at all, see Can I license data to a non-U.S. lab?.

The authorised representative must be in place before market placement

Providers established outside the EU must appoint, by written mandate, an authorised representative established in the Union before placing a GPAI model on the market [2]. Under Article 54, the mandate should let the representative verify that the Annex XI technical documentation exists and that Article 53 obligations, including the copyright policy, have been met. Article 54 also sets how long the representative keeps that documentation (10 years after placement in the published text) and requires it to cooperate with the AI Office; confirm the current wording against the Official Journal.

Authorities may address the representative instead of, or in addition to, the provider. That makes the representative's file your first line of defense. If the copyright policy, opt-out compliance evidence and training data register are not in that file, the representative cannot answer the AI Office without going back to your US team. Article 54 also contains an exception for certain open-source models; check its conditions with counsel rather than assuming it applies.

Training data and location register for EU-bound models

A register that records data provenance and compute location answers both the AI Act policy question and the territorial infringement question. Keep one row per dataset per training run, and give your authorised representative read access.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy it matters
dataset_idDS-2026-0147Joins to license and summary records
source_typeLicensed, supplier-providedTemplate category for the 53(1)(d) summary [5]
license_id / territoryLIC-0098 / worldwideConfirms EU deployment is within the grant
field_of_usePre-training; fine-tuning; commercial deploymentAvoids implied-use disputes
tdm_reservation_checkNot applicable: express licenseEvidence for the 53(1)(c) policy [2]
crawl_componentNoneIf present, record robots.txt handling and crawl dates
training_run_idPT-run-22Links data to a specific model version
compute_provider / regionCloud provider / us-east (all nodes)Territorial evidence in infringement claims [7]
run_dates2026-03-02 to 2026-04-18Matches logs to placement dates
model_version_placed_eumodel-v3.1, placed 2026-06-01Starts the Article 54 documentation retention clock
rep_file_referenceAR-mandate-01 / folder 3.1Shows the authorised representative holds it

Decision table: which duty applies to your situation

The trigger for the EU copyright policy is placing a GPAI model on the EU market, not where you train or where your data came from.

Illustrative example: invented to show structure; it does not describe an available dataset.

Situation53(1)(c) copyright policyAuthorised representativeMain evidence gap
US lab trains in US, offers API to EU customersApplies [1][2]Required [2]Opt-out handling for crawled data
US lab trains in US, releases open weights downloadable in EUApplies [2]Check Article 54 exception [2]Policy often missing for "research" releases
Asia-based lab fine-tunes a third-party model and sells in EUDepends on whether you become the provider of a modified modelRequired if you are the provider [2]Fine-tune data license territory
US lab, model never offered in EUNot triggered on this readingNot requiredGeo-restriction and contractual controls

When modification makes you the provider, our guide on taking on provider duties when fine-tuning with acquired data walks through the analysis. Start from the compliance hub for the full map of AI Act, US state and standards obligations, and the EU AI Act overview for how the regulation affects licensing.

What to ask a data supplier before EU deployment

Ask suppliers questions that produce records, not reassurances. Useful requests include: who owns the records and what consents cover them; whether any portion was collected by crawling, and if so how Article 4(3) reservations were checked; what territory and field of use the license grants; and whether the supplier will help answer AI Office or representative requests [2].

SourceX works this way for operational data. It sources datasets from US companies on request, including support and sales histories, engineering records, documents, and finance and legal workflows, and it serves AI teams wherever they are based. Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery, with diligence materials on source, rights, preparation and allowed use prepared per dataset. Buyers can describe the data their EU-bound model needs.

Sourcing licensed data for a model you will place in the EU

SourceX looks for US businesses that hold the data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is transacted; every release is approved by the supplying company. It does not source scraped web content, and a request does not guarantee a match. Describe your dataset requirements to SourceX.

Sources

  1. Alexander Peukert, Goethe University Frankfurt, "AI Act Copyright (Cracow lecture)" (2026). https://www.jura.uni-frankfurt.de/185651045/2026_04_23_Peukert_Cracow_AI_Act_Copyright.pdf
  2. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models" (2024). https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  3. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
  4. European Commission (AI Office), "The General-Purpose AI Code of Practice" (2025). https://digital-strategy.ec.europa.eu/en/policies/gpai-code-practice
  5. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  6. Alexander Peukert, Goethe University Frankfurt, "Copyright and AI (Tokyo lecture)" (2026). https://www.jura.uni-frankfurt.de/184651156/2026_04_01_Peukert_Tokyo_Copyright_AI.pdf
  7. Osborne Clarke, "Opted out: UK government moves away from preferred position on AI and copyright following report" (2026). https://www.osborneclarke.com/insights/opted-out-uk-government-moves-away-preferred-position-ai-and-copyright-following-report
  8. Akin Gump Strauss Hauer & Feld LLP, "Second District Court Rules AI Training Can Be Fair Use (Kadrey v. Meta)" (2025). https://www.akingump.com/en/insights/ai-law-and-regulation-tracker/second-district-court-rules-ai-training-can-be-fair-use
  9. A&O Shearman, "EU Artificial Intelligence Office publishes the final version of the GPAI Code of Practice" (2025). https://www.aoshearman.com/en/insights/ao-shearman-on-data/eu-artificial-intelligence-office-publishes-the-final-version-of-the-gpai-code-of-practice

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data