Skip to content

Regulation and governance for data buyers

How Long to Keep Training Data and Its Records: Regulatory Retention vs License Deletion Duties

Quick answer

Keep the records, not necessarily the data. Regulators mostly require you to retain documentation about training data: under the EU AI Act, high-risk providers keep technical documentation for 10 years [2], logs for at least six months [2], and Colorado commentary reports a three-year record rule under SB26-189 [5]. Data licenses usually require deleting the raw records at term end. Reconcile the two by keeping manifests, hashes, license copies and deletion certificates, and by negotiating a narrow documentation-retention carve-out.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Which regulations set a retention clock for training data records?

Most retention clocks attach to documentation and logs, and very few require you to keep the training records themselves. The figures below are as of October 2026; check each against current text before you set a schedule.

  • EU AI Act, Article 18 (high-risk providers). Keep the Article 11 technical documentation, quality management system documentation, notified body decisions and the EU declaration of conformity for 10 years after the system is placed on the market or put into service [2]. Annex IV technical documentation includes the description of training, validation and test data, so your dataset descriptions fall inside that 10-year window.
  • EU AI Act, Article 19 (logs). Providers keep automatically generated logs under their control for at least six months, unless applicable Union or national law, in particular data protection law, provides otherwise [2]; Article 26(6) sets the same minimum for deployers. These are runtime logs, not training data, but deployment teams often store them next to evaluation sets.
  • EU AI Act, Article 53 (general-purpose models). Providers must draw up and keep up to date technical documentation, maintain a copyright policy and publish a training content summary [1]. Article 53 sets no fixed retention period, so the practical rule is to keep the documentation for as long as the model stays on the market and you may receive an AI Office request.
  • Timing changes. Regulation (EU) 2026/1744 amended the AI Act and was published on 24 July 2026 [3]. Commentary reports that high-risk dates moved to 2 December 2027 (Annex III) and 2 August 2028 (Annex I), so confirm the date that applies to your system before you start the Article 18 clock.
  • Colorado SB26-189. From January 1, 2027, developers of automated decision-making technology that materially influences consequential decisions must give deployers technical documentation, including categories of training data [4]. Commentary reports a duty to keep compliance records for at least three years [5]. Check the enrolled text before you rely on that number.
  • HIPAA. Privacy Rule documentation under 45 CFR 164.530(j), and Security Rule documentation under 45 CFR 164.316(b)(2), is kept for 6 years from creation or the date last in effect, whichever is later. For health training data, apply that period to de-identification determinations and to data use agreements for limited data sets [6].
  • GDPR. Article 5(1)(e) storage limitation runs the other way: personal data is kept no longer than necessary for its purpose [7]. Data that is truly anonymous under Recital 26 falls outside the GDPR [7], which is why the de-identification record matters as much as the dataset.

For the documentation content itself, see Annex IV technical documentation for datasets and the Colorado SB 26-189 documentation guide.

What must you keep, and what may your license force you to delete?

Split every licensed dataset into two classes of artifact: evidence that proves what you trained on, and the licensed content itself. Regulators ask for the first. Licensors usually control the second.

Evidence artifacts you should expect to keep include the executed license and amendments, the data register entry, file manifests with SHA-256 checksums and record counts, the datasheet or Croissant-RAI metadata file [9], de-identification method records, train/validation/test split definitions, and deletion certificates. None of these needs the underlying records to be useful, and most are small enough to hold for a decade.

Licensed content includes the raw files, working copies in feature stores and vector indexes, cached tokenized shards, and backups. Typical license language requires destruction of all of it at expiry or termination, with a written certification. Derived artifacts sit in between: embeddings, fine-tuned weights and synthetic data generated from the source. Read the license definition of "Derived Data" closely, because the FTC's Everalbum order shows that models built from improperly retained data can themselves be ordered deleted [8]. The algorithmic disgorgement guide covers that risk in depth.

Retention schedule by artifact and regime

A workable schedule assigns each artifact one owner, one trigger event and the longest applicable period, then subtracts anything the license forbids.

Illustrative example: invented to show structure; it does not describe an available dataset.

ArtifactRegulatory driverTrigger eventRetention periodLicense conflict?Disposition
Executed license, amendments, approvalsContract limitation periods; AI Act Art. 18 if part of technical documentation [2]Later of license end or model withdrawal10 years for high-risk EU systems; otherwise counsel-setNone expectedArchive, read-only
Annex IV dataset description, datasheet, Croissant-RAI file [9]AI Act Art. 18 [2]; Colorado records (reported 3 years) [5]System placed on market10 years (EU high-risk)Low: contains descriptions, not recordsArchive with model version
Manifest: file paths, SHA-256, row counts, split IDsSupports Art. 18 documentation and auditsDataset frozen for trainingSame as model documentationLow, provided no record content is embeddedArchive
De-identification determination, method log, QA sample resultsHIPAA 6 years [6]Later of creation or last in effect6 years minimumSample excerpts may count as licensed contentKeep report; delete excerpts
Raw licensed records and working copiesGDPR storage limitation if personal data [7]License expiry or terminationLicense term onlyYes: deletion clause governsDelete and certify
Embeddings, vector indexes, tokenized shardsNone specificLicense expiryPer "Derived Data" clauseOften yesDelete unless carved out
Fine-tuned weightsModel documentation duties [2][1]Model retirementModel life plus documentation periodDepends on licenseRetain only if license permits
Inference logs of high-risk systemsAI Act Art. 19 [2]Log creationAt least 6 monthsRarelyRotate
Deletion certificateEvidence of license complianceDeletion completedLongest of the aboveNoneArchive permanently with register

How do you reconcile a deletion clause with a 10-year documentation duty?

Negotiate the conflict away before signature, because a deletion clause signed without a carve-out leaves you choosing between breach and noncompliance. Three clauses do most of the work.

First, a documentation retention carve-out: the licensee may keep the license, manifests, checksums, metadata, evaluation reports and de-identification records after termination, for as long as law requires, provided they contain no licensed record content. Second, a regulatory access clause: if a market surveillance authority, notified body or the AI Office requests the training data, the licensee may retain or re-obtain a frozen snapshot under confidentiality, or the licensor agrees to cooperate. The regulator access guide sets out who can ask for what. Third, a reproducible snapshot escrow: a hashed, encrypted copy held by the licensor or a neutral party, so you can prove that a given checksum matches the content you trained on without holding the content yourself.

If the licensor refuses all three, record the gap in your risk register and make sure your manifest is detailed enough to stand in for the data. A manifest that lists source system, export date, field schema, record count and per-file SHA-256 lets an auditor confirm identity against the licensor's copy. If you are scoping a new purchase, you can put these clause requirements in your request to SourceX for buyers. Licensing terms on this point are covered in retention and deletion in data licensing.

A litigation or regulatory hold suspends deletion for every artifact it covers, including raw licensed data your license says to destroy. In US litigation, Federal Rule of Civil Procedure 37(e) allows sanctions when electronically stored information that should have been preserved is lost, and deleting a training corpus on schedule after a hold notice is a classic failure mode.

Build the suspension into the process rather than relying on memory:

  • Tag each dataset in the register with a hold_status field and a hold_ref pointing to the hold notice.
  • Make the deletion job read hold_status and refuse to run on any held dataset, including backups and vector stores.
  • Notify the licensor in writing that a hold prevents deletion, citing the clause that permits retention where law requires it.
  • When the hold lifts, run deletion within the period the license allows and issue the certificate then.

Copyright disputes over training data, such as the cases tracked on the compliance hub, are exactly where holds arise. Expect counsel to ask for manifests and license copies first.

Running deletion so the certificate holds up

A deletion certificate is only as credible as the inventory behind it. Before term end, list every location that holds the dataset: object storage buckets, training cluster scratch disks, feature stores, vector databases, notebooks, evaluation harnesses and backup snapshots with their own retention cycles.

Illustrative example: invented to show structure; it does not describe an available dataset.

deletion_certificate:
  dataset_id: DS-2026-0412
  license_ref: LIC-0412-v2
  trigger: license_expiry
  locations_purged:
    - s3://training-raw/ds-0412/        # versioning disabled, noncurrent versions removed
    - vector-index: support-rag-prod     # namespace dropped, rebuild verified
    - backups: nightly snapshots         # aged out after 35-day cycle
  retained_under_carve_out:
    - manifest_sha256.csv
    - croissant_rai.json
    - deid_method_report.pdf
  holds_checked: none_active
  verified_by: records-manager
  completed: 2026-09-30

Record backup expiry honestly. If snapshots age out on a 35-day cycle, the certificate should say deletion completes when the last snapshot expires, not on the day the bucket was emptied. The related question of whether labs delete after training is answered in do AI labs delete data after training.

Setting periods for evaluation and RAG data

Evaluation and retrieval sets follow the same split but usually need longer live retention, because they stay in use after training ends. A held-out test set may need to persist through every model version you report results for, which can outlast a training-data license term.

Price that into the license at the start: ask for a term that covers the evaluation life of the model, or a separate right to retain the test split. RAG corpora are worse, because the content is served at inference time; once the license ends, the index must go, and answers that depend on it change. The validation and test data requirements page covers what regulators expect of those sets, and how long buyers keep licensed data gives typical patterns.

Plan retention before you license the data

SourceX sources operational datasets from US companies on request and manages the licensing process; every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Retention and deletion needs belong in the request from the start, since pricing and allowed uses are agreed per deal and nothing is contracted until a supplier agrees. Describe the data and your documentation needs at SourceX for buyers.

Frequently asked questions

Does the EU AI Act require keeping the training data itself for 10 years?

Article 18 lists documentation, not the dataset [2]. The Annex IV description of training data falls within it, so the description is kept for 10 years. Whether authorities can also demand the data is a separate access question, which is why a regulatory access clause matters.

Is Colorado's three-year rule confirmed?

Commentary reports it [5], and the official bill summary confirms the developer documentation duty from January 1, 2027 [4]. Check the enrolled text of SB26-189 before you set the period.

Can we keep embeddings after the license ends?

Only if the license definition of derived data permits it. Treat embeddings as licensed content unless a clause says otherwise, because they can encode source records.

Sources

  1. European Commission, AI Act Service Desk, "Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  2. Official Journal of the European Union, via EUR-Lex, "Regulation (EU) 2024/1689 (Artificial Intelligence Act)" (2024). https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
  3. Official Journal of the European Union, via EUR-Lex, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
  4. Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
  5. Promise Legal, "Colorado AI Act Compliance: SB 26-189 Developer & Deployer Guide" (2026). https://blog.promise.legal/colorado-ai-act-compliance-guide/
  6. Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - De-identification and limited data sets". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  7. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  8. Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
  9. Jain et al., MLCommons (arXiv:2407.16883), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data