Skip to content

Provenance, rights and permitted use

Matching Each Record to the Notice and Terms in Force When It Was Collected

Quick answer

The notice that governs a record is generally the privacy policy and terms version in force when that record was collected, not the version live today. To document this for training data, collect every notice and terms version with its effective date, join each record's collection timestamp to that timeline, store the resulting version ID on the record, and set permitted uses per cohort. Records collected before AI training was disclosed need their own decision, because a later policy change does not automatically reach them.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why the collection-time notice controls old records

The collection-time notice controls because it is the promise the person relied on when they handed over the data. The FTC's Office of Technology said in February 2024 that a company adopting more permissive practices, such as using data for AI training, through surreptitious or retroactive changes to its terms or privacy policy could be engaging in unfair or deceptive conduct [1]. A month earlier, the same office warned that companies may be liable under FTC-enforced laws when they break promises not to use customer data for undisclosed purposes such as model training [2]. Both are staff blog posts rather than rules, but they describe a theory buyers should plan around as of October 2026.

In the EU, the same question surfaces through purpose limitation and reasonable expectations. EDPB Opinion 28/2024 treats what data subjects could reasonably expect at the time and in the context of collection as part of any legitimate-interest assessment for AI model development [4]. A notice that said nothing about training is weak evidence that training was expected.

Sector rules add hard limits that a policy rewrite cannot loosen. Under Regulation P, a recipient of nonpublic personal information from a nonaffiliated financial institution is limited in how it may reuse and redisclose it [6]. Under HIPAA, a business associate may use protected health information only as its business associate contract permits or as the law requires [7]. Those constraints travel with the record from the day it was collected.

What changes between notice versions, and why buyers care

The changes that matter are the clauses that define purpose, sharing and retention, and they change more often than most teams assume. An audit of web data permissions by the Data Provenance Initiative found that terms of service and robots.txt restrictions shifted over short periods and frequently contradicted each other [3]. Company privacy notices and customer terms behave the same way: product launches, acquisitions and new state laws each trigger a rewrite.

For a training-data buyer, the clauses to diff across versions are these:

  • Purpose statements: "to provide and improve our services" versus explicit language about developing or training machine-learning models.
  • Sharing and sale: whether disclosure to third parties for their own purposes, including licensing, was ever described.
  • De-identified data: whether the notice reserves the right to use or share de-identified or aggregated data, and how it defines that term.
  • Retention: whether records past a stated retention period should still exist at all.
  • Opt-out mechanics: whether users were offered a training or sharing opt-out, and where those choices are stored.
  • Customer contract overrides: B2B master agreements and DPAs often restrict use more tightly than the public policy; see customer contracts and DPAs for training use.

The owner pages on whether licensing data violates a privacy policy and whether a policy update is needed before licensing cover the supplier's side of this question. This page covers the mapping method a buyer's engineers and governance lead run on the delivered records.

Building the notice version register

Build the register first, because every later join depends on accurate effective dates. Each row is one version of one governing document, and each document type gets its own timeline.

Where to find historical versions: the supplier's legal or compliance team usually keeps a policy archive or a document-management history; CMS revision logs and git history for the website repository record publish dates; app-store listings carry their own privacy-policy links; and B2B terms often live in a CLM system such as the supplier's contract repository. Archived public snapshots can corroborate a date but should not be the only evidence.

Illustrative example: invented to show structure; it does not describe an available dataset.

version_iddocumenteffective_fromeffective_totraining_disclosedthird_party_licensing_discloseddeid_use_reservedevidence
PP-2016-03Consumer privacy policy2016-03-012019-12-31nononoArchived PDF, CMS log
PP-2020-01Consumer privacy policy2020-01-012023-06-14nonoyesCMS log, counsel memo
PP-2023-06Consumer privacy policy2023-06-15openyesyesyesGit commit, email notice log
TOS-2018-09Terms of service2018-09-102024-02-28nonon/aCLM export
TOS-2024-03Terms of service2024-03-01openyesnon/aCLM export, in-app consent log

Three details prevent the most common errors. Record effective dates in UTC with an explicit timezone, because a notice published late on the last day of a month in one region is already the next day in another. Keep effective_to as the day before the next version, so intervals never overlap. Record how users were notified (email, in-app banner, click-through re-acceptance), because a silent web update and an affirmative re-acceptance carry very different weight [1].

Joining records to the version in force

The join is an interval lookup: for each record, find the notice version whose effective window contains the record's collection timestamp. In SQL this is a range join on collected_at >= effective_from AND collected_at < effective_to_exclusive; in pandas, merge_asof on sorted timestamps does the same job for one document type at a time.

The hard part is choosing the right timestamp. Use the moment the data was provided by the person, not when it was later loaded, migrated or exported. Common candidates:

  • Support tickets: the ticket or message created_at, per message rather than per ticket, because long threads can span a policy change.
  • CRM records: the activity or note creation date, not the account LastModifiedDate, which updates on unrelated edits.
  • Account-level data: the signup date for data captured at signup, plus the date of each later submission.
  • Migrated systems: original source-system timestamps where they survived; an import date is not a collection date.

Records with no reliable timestamp should default to the most restrictive version in the timeline, and the default should be flagged rather than hidden. Where a user re-accepted later terms, store both the collection-time version and the latest accepted version; whether re-acceptance extends to earlier data is a legal call, not an engineering one.

Per-record fields to store

Store the version ID on every record so the decision survives splits, sampling and deduplication. These fields slot into a broader permitted-use metadata schema and support record-level provenance when dataset-level documentation is too coarse.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "tkt-000184-msg-03",
  "collected_at": "2021-11-04T17:22:09Z",
  "collected_at_source": "helpdesk.message.created_at",
  "notice_version_id": "PP-2020-01",
  "terms_version_id": "TOS-2018-09",
  "contract_ref": null,
  "version_assignment": "timestamp_join",
  "training_disclosed_at_collection": false,
  "opt_out_status": "none_recorded",
  "cohort": "pre_disclosure",
  "permitted_uses": ["eval_internal"],
  "deidentification_method_ref": "deid-run-2026-09-12"
}

version_assignment should take values such as timestamp_join, default_most_restrictive or manual_review, so auditors can see which records were inferred. Keep the register itself versioned too; if a new historical policy turns up, rerun the join and diff the cohort assignments.

Handling records collected before AI training was disclosed

Pre-disclosure records are the retroactivity problem, and the safe default is to treat them as a separate cohort with narrower permitted uses until counsel decides otherwise. FTC staff focused their 2024 warning on applying newly permissive practices to data collected under older promises [1].

Illustrative example: invented to show structure; it does not describe an available dataset.

CohortTypical evidenceOptions buyers see in practiceQuestions for counsel
Post-disclosure, affirmative consentClick-through or opt-in log tied to user IDPre-training, SFT, eval as licensedDoes consent cover licensing to a third party, or only the supplier's own models?
Post-disclosure, notice onlyPolicy update plus email noticeSFT and eval after de-identificationIs notice alone enough for the purpose and jurisdiction?
Pre-disclosure, de-identifiedDeid method record, sample checkNarrow uses, often eval onlyDoes the original notice's de-identified-data clause reach this use?
Pre-disclosure, identifiableCollection-time notice silent on trainingExclude, or seek fresh consentIs any lawful basis available without new consent?
Sector-restricted (GLBA, HIPAA)Contract or statute governsOnly what the contract or rule permits [6] [7]Does the BAA or Regulation P exception allow de-identification and onward use?

De-identification changes the analysis but does not erase it: an older notice may still have promised not to share data for unrelated purposes, and a buyer's model documentation will record where the records came from. The training data use register is the natural place to carry each cohort's permitted uses forward into model cards and internal approvals. For the privacy controls themselves, see the de-identified training data guide.

Documentation regulators and model cards now expect

Collection dates are becoming a disclosure item, so the mapping does double duty. California's AB 2013, whose posting deadline was 1 January 2026, requires developers of generative AI systems offered to Californians to publish training-data documentation, including whether datasets contain personal information and the time period during which the data was collected [5]. A per-record notice version makes that time period, and the terms that applied during it, straightforward to state accurately.

For EU-facing models, the reasonable-expectations reasoning in EDPB Opinion 28/2024 means a controller should be able to show what people were told when their data was gathered [4]. The register plus per-record version IDs is that evidence.

Checklist: what to request from a supplier

Ask for these items before accepting records collected over many years; the companion page on consent and notice records covers how to verify each record type, and the provenance buyer's guide places this check within the full rights review. Teams that want SourceX to look for US businesses holding records like these can send a data request through the buyers page; a request does not guarantee a match.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Every privacy notice and terms version in scope, with effective dates and evidence of each date.
  • Notification method per change (silent update, email, in-app re-acceptance) and the logs that prove it.
  • Relevant B2B contracts, DPAs or BAAs, and which customer accounts each governs.
  • The source field used for each record's collection timestamp, and how migrated records were handled.
  • Opt-out and deletion request logs, with dates, so withdrawn records can be removed.
  • A cohort summary: record counts before and after the first training or licensing disclosure.
  • The de-identification method record and sample-check results for each cohort.

Sourcing customer records with a documented notice history

SourceX sources operational datasets, such as support and sales histories, from US businesses that hold the data described and manages licensing and ongoing purchases. Every dataset is reviewed for ownership and consents and delivered under a license that sets records, uses, term and delivery, and the supplying company approves each release. If you need records with a notice history you can map, describe the data you need to SourceX's buyer team.

Frequently asked questions

Which privacy policy applies to data collected years ago?

Generally the version in force at collection, unless the person later gave affirmative consent that clearly covers earlier data. A silently updated policy is weak support for new uses of old data [1].

Can a company change its terms of service to allow AI training on existing data?

It can change terms going forward, but FTC staff have warned that applying more permissive practices retroactively, without clear notice and consent, may be unfair or deceptive [1] [2]. Sector contracts and statutes may forbid it outright [6] [7].

What if records have no reliable collection date?

Assign the most restrictive version in the timeline, mark the assignment as a default, and exclude those records from uses that depend on a later disclosure.

Sources

  1. Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  2. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  3. arXiv (Data Provenance Initiative), "Consent in Crisis: The Rapid Decline of the AI Data Commons" (2024). https://arxiv.org/pdf/2407.14933
  4. European Data Protection Board, "EDPB opinion on AI models: GDPR principles support responsible AI" (2024). https://www.edpb.europa.eu/news/edpb-opinion-on-ai-models-gdpr-principles-support-responsible-ai_en
  5. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  6. Consumer Financial Protection Bureau, "12 CFR 1016.11 Limits on redisclosure and reuse of information". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
  7. U.S. Department of Health and Human Services, "Business Associate Contracts". https://www.hhs.gov/hipaa/for-professionals/covered-entities/sample-business-associate-agreement-provisions/index.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data