Regulation and governance for data buyers
Fine-Tuning With Acquired Data: When You Take On Provider Duties Under the EU AI Act or Developer Duties Under AB 2013
Quick answer
Fine-tuning a third-party model rarely makes you an EU general-purpose AI (GPAI) model provider: the European Commission's July 2025 GPAI guidelines reportedly use an indicative test of modification compute above one third of the original model's training compute. California AB 2013 is different. Fine-tuning that materially changes a public-facing generative system's functionality or performance can make you a "developer" who must post training-data documentation before release [7][8]. Either way, keep a per-version register of every acquired fine-tuning dataset.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
The EU test: one third of original training compute
Under the EU AI Act, a downstream modifier becomes the provider of a GPAI model only when the modification is substantial, and the Commission's July 2025 guidelines use compute as the indicative proxy. Law-firm commentary on the guidelines reports that a modifier is presumed to become the provider when the compute used for the modification exceeds one third of the compute used to train the original model. The guidelines describe this as an indicative criterion, read together with whether the change significantly alters the model's generality, capabilities or systemic risk.
When the original compute is not disclosed and cannot be estimated, summaries report fallback thresholds expressed as one third of the GPAI and systemic-risk compute levels (10^23 and 10^25 FLOP); secondary sources differ on the exact fallback, so check the official text. If you modify a model already classed as systemic-risk enough to become its provider, the modified model is reportedly presumed systemic-risk as well.
In practice, a LoRA or full-parameter supervised fine-tune (SFT) on tens of thousands of support tickets, or a DPO run on a few thousand preference pairs, uses a tiny fraction of a frontier base model's pretraining compute. Continued pretraining on hundreds of billions of domain tokens against a smaller open-weight base is where the ratio can approach the threshold. Do the arithmetic and file it: approximate FLOP as 6 x parameters x training tokens for each run, and compare it against the base model's disclosed or estimated figure.
What a modifier-provider must document
If you cross the threshold, your Article 53 obligations reportedly cover the modification, not the whole base model, unless the modified model carries systemic risk. Article 53 requires providers to keep technical documentation, give information to downstream integrators, maintain a copyright-compliance policy that honors text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content [1]. For a modifier, that means documenting your fine-tuning and continued-pretraining corpora, not re-describing the base model's pretraining data.
The public summary must follow the AI Office template published on 24 July 2025 [2]. Expect it to ask for data modalities, size bands, main source categories (for example licensed private data from third parties), and how you handled rights reservations; read the field list in the template files rather than relying on summaries [2]. The voluntary GPAI Code of Practice (10 July 2025) offers a documentation form and a copyright chapter that many providers use as the route to show compliance [3]. See Article 53 training data obligations and completing the GPAI documentation form.
Timing matters as of October 2026. GPAI duties have applied since 2 August 2025, and the AI Office's enforcement powers began in August 2026 for models placed on the market after the start date [4]. A fine-tuned model you release in the EU today, if it crosses the threshold, is a newly placed model.
Below the threshold: Article 10 can still apply
Staying under one third of compute does not exempt you from data duties if the fine-tuned model sits inside a high-risk AI system. Article 10 requires training, validation and test sets for high-risk systems to meet governance and quality criteria: documented collection processes and origin, preparation steps such as annotation and cleaning, bias examination, and checks that data are relevant, sufficiently representative and as free of errors and complete as possible [5]. That applies to your fine-tuning set because it is training data for the system, whatever the base model's GPAI status.
Regulation (EU) 2026/1744 amended the AI Act and, per secondary commentary, moved high-risk application dates to 2 December 2027 for Annex III uses and 2 August 2028 for Annex I products; Article 10 was also amended [6]. A credit-scoring, hiring, or life- or health-insurance pricing system fine-tuned on acquired loan files or underwriting notes is the typical case. See Article 10 data governance for licensed data and Annex IV dataset descriptions.
AB 2013: fine-tuning as a substantial modification
AB 2013 reaches fine-tuners directly because its "developer" definition includes anyone who substantially modifies a generative AI system or service made available to Californians for public use [7][9]. Coverage of the statute reports that "substantial modification" includes a new version, release or update that materially changes functionality or performance, including through retraining or fine-tuning [8]. Each such release needs posted documentation before it is made available [8].
The documentation in Civil Code section 3111 is a high-level summary of the datasets used in development, including sources or owners, how the data further the system's purpose, approximate data-point counts, data types, whether data are protected by copyright, trademark or patent, whether they were purchased or licensed, whether they include personal information or aggregate consumer information, cleaning or processing steps, collection periods, first-use dates, and synthetic data use [7][9]. The posting deadline was 1 January 2026, and the duty covers systems released since 1 January 2022 [7].
Scope is the open question for enterprises. The statute speaks to systems made available to the public; an internal-only assistant used by employees may fall outside it, but that turns on facts such as contractor or customer access, so flag it for counsel [7]. For a fine-tune, a defensible scope is your fine-tuning datasets plus a reference to the base model developer's own AB 2013 disclosure. See California AB 2013 supplier records and disclosure regimes compared.
Decision table: which regime your fine-tune triggers
The same fine-tuning run can trigger zero, one or all three regimes, so test each separately.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Scenario | EU GPAI provider (Art. 53)? | EU Art. 10 (high-risk)? | AB 2013 developer? |
|---|---|---|---|
| LoRA SFT of a 70B open-weight model on 40k licensed support tickets; internal help-desk copilot | Unlikely: far below one third of compute | Only if used in an Annex III use | Possibly not: internal-only, confirm with counsel |
| Same fine-tune exposed as a customer-facing chat agent in California | Unlikely | Only if high-risk use | Likely, if functionality materially changes |
| Continued pretraining of a 7B base on 600B tokens of engineering records, released on Hugging Face in the EU | Possible: compute the ratio | Depends on downstream use | Likely, if public and generative |
| DPO on underwriting notes for an EU life-insurance pricing system | Unlikely | Yes, from the applicable high-risk date | Only if offered to the California public |
A fine-tuning data register per model version
One register entry per model version lets you answer the EU summary template, an Article 10 file and an AB 2013 posting from the same facts. Record it when you acquire data, not when a regulator asks, because supplier facts such as consent basis and collection dates are hard to reconstruct later.
Illustrative example: invented to show structure; it does not describe an available dataset.
model_version: helpdesk-assist-v3.2
base_model: open-weight-70b-instruct (base developer AB 2013 disclosure: <URL>)
method: LoRA SFT + DPO
modification_compute_flop: 3.1e20 # 6 x params x tokens, method noted
base_training_compute_flop: 6.0e24 # disclosed | estimated (state which)
ratio: 0.00005 # vs 0.333 indicative EU threshold
datasets:
- id: ds-017
description: B2B SaaS support tickets with agent resolutions
owner: licensed from US company (name per license)
acquisition: licensed
records: ~40,000 tickets
collection_period: 2021-01 to 2025-06
personal_info: names/emails/account numbers replaced; method recorded
ip_status: supplier-owned text; license permits fine-tuning
processing: dedup, PII replacement, length filter
first_used: 2026-09-14
synthetic: no
release_targets: [EU, California public]
eu_status: not provider (ratio); Art. 10 n/a (not high-risk)
ab2013_status: developer (public release, material change); posted 2026-10-01
The fields mirror AB 2013's section 3111 list and the Article 10 origin and preparation records [5][7]. Pair the register with license terms that let you keep these records and share them with authorities; see regulator access to licensed datasets and retention versus deletion duties.
What to ask data suppliers before you fine-tune
Supplier records decide whether you can complete any of these disclosures, so request them in diligence. Ask for the data's origin and ownership chain, the consent or notice basis for personal data, the de-identification method and residual-risk notes, collection date ranges, approximate record counts, whether any content is synthetic or model-generated, and whether the license permits fine-tuning, record retention and regulator disclosure. For scope language, see fine-tuning-only data licenses; for quality checks, see evaluating a fine-tuning dataset before buying.
SourceX sources operational datasets from US companies on request, with each dataset rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Buyers who need that kind of documented provenance for a fine-tuning or supervised fine-tuning program can describe the data they need.
Get documented fine-tuning data for your next model version
SourceX looks for US businesses holding the operational data you describe, and every release is approved by the supplying company. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, and nothing is contracted until a supplier agrees. Explore the compliance hub or start a buyer request.
Frequently asked questions
Does a LoRA adapter count as a modification of the model?
Yes, an adapter changes the model's behavior, but provider status turns on the compute and significance test, not on the technique. Most adapter runs fall far below one third of base training compute. Keep the calculation on file anyway.
If I become an EU provider, must I summarize the base model's pretraining data?
Law-firm commentary on the guidelines reports that a modifier's obligations are limited to the modification. Your summary should describe your fine-tuning and continued-pretraining data and point to the original provider's documentation for the rest [2].
Does an AB 2013 posting need updating for every re-tune?
A re-tune that materially changes functionality or performance is a substantial modification, and documentation is due before that release [8]. Routine refreshes that do not materially change behavior may not qualify; document why. See fine-tuning data refresh cadence.
Sources
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- European Commission (AI Office), "The General-Purpose AI Code of Practice" (2025). https://digital-strategy.ec.europa.eu/en/policies/gpai-code-practice
- Latham & Watkins, "EU AI Act GPAI model obligations in force and final GPAI Code of Practice in place" (2025). https://www.lw.com/en/insights/2025/09/eu-ai-act-gpai-model-obligations-in-force-and-final-gpai-code-of-practice-in-place
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- Official Journal of the European Union (EUR-Lex), "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- Conventus Law, "US: California's AB 2013 Requires Generative AI Data Disclosure By January 1, 2026". https://conventuslaw.com/report/us-californias-ab-2013-requires-generative-ai-data-disclosure-by-january-1-2026/
- Perkins Coie, "AB 2013: New California AI Law Mandates Disclosure of GenAI Training Data". https://perkinscoie.com/insights/update/ab-2013-new-california-ai-law-mandates-disclosure-genai-training-data
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.