Regulation and governance for data buyers
UK AI Training and Copyright in 2026: Why Commercial Training Still Needs a License
Quick answer
As of October 2026, the UK has no text and data mining exception for commercial AI training. Section 29A of the Copyright, Designs and Patents Act 1988 (CDPA) covers copies made for computational analysis for non-commercial research only, by someone with lawful access [3]. The government's 18 March 2026 Report on Copyright and Artificial Intelligence dropped the proposed opt-out exception as its preferred option and committed to no replacement [1][5]. A lab training commercially on UK-protected works therefore needs a license, or a defensible analysis of where training happens.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What section 29A actually permits
Section 29A is a narrow research exception, not a training permission. It allows a person with lawful access to a work to make a copy to carry out computational analysis of anything recorded in the work, where the sole purpose is non-commercial research and the copy carries a sufficient acknowledgement unless that is impossible [3]. The same section makes it an infringement to transfer the copy to anyone else, or to use it for another purpose, without the rightsholder's authorization, and makes contract terms that try to override the exception unenforceable [3].
The Intellectual Property Office's guidance from when the exception was introduced is blunt about the commercial line. Contract research carried out for an outside company is unlikely to qualify as non-commercial, and if the purpose is not solely non-commercial the researcher is very likely to infringe [4]. For counsel, that removes the familiar workaround of running "research" in a university or a research subsidiary and then handing the weights or the corpus to a product team.
Three failure modes recur in diligence:
- Purpose drift. A corpus assembled under s29A at an academic partner is later used to pre-train a model that ships in a paid API. The original copies were made for a non-commercial purpose; the later use is not covered [3][4].
- Onward transfer. A research group shares its s29A copies with a commercial collaborator. Transfer without authorization is itself outside the exception [3].
- No lawful access. The exception presupposes lawful access, so works obtained from pirate mirrors or by breaching paywalls never qualify, whatever the purpose. See lawful access and pirated sources for how acquisition route changes risk.
Database right is a separate layer. A structured dataset can carry sui generis database right alongside copyright in its contents, and s29A is a copyright provision, so check the database-right position on its own terms.
What the March 2026 report changed, and what it did not
The report changed government policy direction but not the law. It was published on 18 March 2026 under the reporting duty in section 136 of the Data (Use and Access) Act 2025, which required a report on the use of copyright works in developing AI systems; that section creates a reporting obligation and amends no exception [1][2].
The December 2024 consultation (CP 1205) had put forward, as the government's preferred option, a TDM exception covering commercial use, subject to rightsholders being able to reserve their works, with transparency measures alongside [8]. The consultation drew more than 11,500 responses, mostly from rightsholders, and commentators report that 81% favored requiring licenses in all cases [5]. The report stepped back from the opt-out model, endorsed no alternative, and signaled that reform is not imminent [5][7].
The stated next steps are to let industry-led licensing develop, to monitor developments in other jurisdictions, and to do more work on transparency and enforcement [6][7]. Linklaters characterizes the result as "hurry up and wait" [7]. For planning purposes, assume s29A in its current form governs any training you run in the UK through at least the next model cycle, and track the dated sources rather than predicting a statute.
Note that the UK's 2025 data reforms did not touch this point: most Data (Use and Access) Act changes were in force by 5 February 2026, and s29A remains non-commercial only. Data protection obligations for the same corpora are covered on the UK GDPR and AI data licensing page.
Where training happens: the territorial question
UK copyright infringement is territorial, so the place where copies are made during training matters as much as the content. Primary infringement under the CDPA requires a restricted act, such as copying, done in the UK. In the most-watched UK image-model litigation, the claimant's primary training claims were dropped during trial, and the secondary-infringement ruling is now on appeal [6].
That outcome is not a safe harbor. It turned on evidence about one developer's compute location; secondary infringement (importing or dealing in an infringing article) rests on different facts and law and is the subject of the pending appeal, and output-side claims remain untested in UK disputes as of October 2026. Location also tends to change over a model's life: a base model trained abroad may be fine-tuned, evaluated or continually updated on UK infrastructure, and each of those runs makes fresh copies.
For counsel, the practical consequence is evidential. If your position depends on training location, you need records that prove it, per run, not a policy statement:
- Compute region and provider for each pre-training, fine-tuning and evaluation run, with job IDs and dates.
- Where the training corpus was stored, cached and pre-processed (tokenization and deduplication pipelines make copies too).
- Which teams in which countries had access to raw copies.
Location analysis also interacts with other regimes. A model trained in the US and placed on the EU market still brings the EU copyright-policy duty under AI Act Article 53(1)(c), which requires honoring Article 4(3) DSM reservations [9]. See EU copyright rules for models trained outside the EU and the country-by-country TDM comparison.
Decision table: which UK route fits your training use
The route depends on purpose, location and who holds the copies. Use this as a first-pass triage before counsel's full analysis.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Scenario | UK copies made? | s29A available? | Practical route |
|---|---|---|---|
| University lab trains a model for a published paper, no commercial sponsor | Yes | Possibly, if purpose is solely non-commercial and access is lawful [3][4] | Document purpose, acknowledgements and access; bar onward transfer |
| Same lab, sponsored by a company that receives the weights | Yes | Unlikely; contract research for a company is probably commercial [4] | License the works or restructure |
| UK-based product team fine-tunes on licensed customer-support transcripts | Yes | Not needed | License that names fine-tuning, the model scope and UK processing |
| Pre-training run entirely on non-UK compute, model later served in the UK | Disputed | Not applicable | Location evidence pack plus licenses for high-value sources |
| Continual fine-tuning in a UK region on a scraped web crawl | Yes | No | Remove unlicensed material or obtain licenses; no opt-out exception exists |
| Evaluation set built from paywalled articles via a UK subscription | Yes | No (commercial purpose; subscription terms likely restrict) | License that covers evaluation use explicitly |
What a UK-ready training license should state
A license that will hold up in the UK needs to authorize the specific acts training involves, because no exception fills the gaps. The clauses that matter most for counsel are the ones that map to the copies your pipeline actually makes.
Illustrative example: invented to show structure; it does not describe an available dataset.
UK training-rights checklist for a licensed dataset
- Licensed acts. Reproduction for ingestion, pre-processing (tokenization, deduplication, filtering), training, fine-tuning and evaluation, named separately.
- Territory. Whether copies may be made in the UK and in which other regions; avoid a license silent on where compute runs.
- Model scope. Which models, versions and derivatives may be trained, and whether weights may be distributed or served.
- Record definition. Exact record types, fields and date ranges, so the license maps to the manifest you ingest.
- Chain of title. Who owns copyright and any database right in the records, and how employee- or contractor-authored content was assigned. See copyright in operational business records.
- Personal data. De-identification method, residual-risk statement and UK GDPR basis, kept separate from the copyright grant.
- Term and post-term. What happens to stored copies and trained weights when the term ends.
- Transparency inputs. Source descriptions you can reuse for future UK transparency measures and existing EU and California disclosures; see disclosure requirements compared.
Keep the signed license, the delivered manifest and the run-level location log together. If the UK later introduces transparency or enforcement measures, as the report signals it will study, those three records are what you will be asked for [6][7].
How UK counsel should source commercial training data now
Because the government is leaving licensing to the market, the commercial default is direct licensing from the parties that hold the data [6]. For web-scale text, that often means publisher and collective deals. For domain data, such as support histories, engineering tickets or finance workflows, the rightsholder is usually the business that generated the records, which simplifies chain of title compared with crawled content.
SourceX sources operational datasets from US companies on request and manages the licensing process; datasets are not held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, permitted uses, term and delivery, and each release is approved by the supplying company. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked, though no method is perfect. Teams that want to scope a request can describe the data they need through the SourceX buyer intake.
For the wider UK legal picture, including data protection and sector rules, see the United Kingdom country page and the compliance hub for data buyers.
Get licensed training data for UK model development
SourceX finds US businesses that hold the operational data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Nothing is contracted until a supplier agrees, and delivery runs through private, access-controlled workflows after an executed agreement. Describe your dataset at sourcex.si/buyers.
Frequently asked questions
Does the UK have an opt-out TDM exception like the EU?
No. The EU's DSM Article 4 exception permits commercial mining unless rights are reserved; the UK consulted on a similar model in 2024 but stepped back from it in March 2026 [5][8]. Compare the EU mechanics on the DSM Article 4 opt-out page.
Can a contract override s29A?
Not to block it: s29A makes terms that purport to prevent or restrict copying under the exception unenforceable [3]. But because s29A only covers non-commercial research, that protection rarely helps a commercial trainer.
Did the 2026 report reject mandatory licensing too?
The report endorsed no alternative model; it chose to monitor, gather evidence and let industry-led licensing develop [6][7]. Treat the question as open as of October 2026.
Sources
- UK Government (GOV.UK), "Report on Copyright and Artificial Intelligence" (2026). https://www.gov.uk/government/publications/report-and-impact-assessment-on-copyright-and-artificial-intelligence/report-on-copyright-and-artificial-intelligence
- legislation.gov.uk, "Data (Use and Access) Act 2025, section 136" (2025). https://www.legislation.gov.uk/ukpga/2025/18/section/136
- UK Intellectual Property Office, "Copyright, Designs and Patents Act 1988 - Consolidated (section 29A)". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
- UK Intellectual Property Office, "Exceptions to copyright: Research" (2014). https://assets.publishing.service.gov.uk/government/uploads/system/uploads/attachment_data/file/375954/Research.pdf
- Reed Smith, "UK Copyright and AI Report: The Opt-Out Is Dead, but What Comes Next?" (2026). https://www.reedsmith.com/articles/uk-copyright-and-ai-report-the-opt-out-is-dead-but-what-comes-next/
- Osborne Clarke, "Opted out: UK government moves away from preferred position on AI and copyright following report" (2026). https://www.osborneclarke.com/insights/opted-out-uk-government-moves-away-preferred-position-ai-and-copyright-following-report
- Linklaters, "Hurry up and wait: UK government still non-committal on AI copyright reforms" (2026). https://www.linklaters.com/en/insights/blogs/digilinks/2026/march/hurry-up-and-wait-uk-government-still-non-committal-on-ai-copyright-reforms
- UK Government, "Copyright and AI: Consultation CP 1205" (2024). https://assets.publishing.service.gov.uk/media/6762c95e3229e84d9bbde7a3/241212_AI_and_Copyright_Consultation_print.pdf
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.