Skip to content

Regulation and governance for data buyers

Algorithmic Disgorgement: When Improperly Obtained Data Forces Model Deletion

Quick answer

Algorithmic disgorgement is a remedy in which the Federal Trade Commission requires a company to delete not only improperly obtained data but also the models, algorithms and derived artifacts built from it. The FTC has used it in settled orders since at least 2021, typically where data was collected deceptively or used contrary to privacy promises [1][5]. For AI buyers, the practical defense is acquisition hygiene: verify the supplier's notices and consents, obtain written representations, and keep lineage that shows which models touched which dataset.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What the remedy actually requires

The remedy reaches the work product, not just the raw records. In the 2021 photo-storage matter, the FTC alleged that an app developer misrepresented its use of facial recognition, and the order required deletion of photos and videos from deactivated accounts plus the face embeddings and any models or algorithms developed with users' images [1][3]. The order was finalized in May 2021 [2]. A statement by then-Commissioner Rohit Chopra framed the point bluntly: a company should not keep the fruits of data it obtained deceptively [4].

By 2024, FTC technologists listed several actions requiring firms to delete algorithms trained on illegally collected data, naming matters involving Rite Aid, Amazon Ring and Avast [5]. In the facial recognition retail matter, the agency also alleged deployment without adequate testing that produced false matches, and the proposed order included a five-year ban on the technology for surveillance [6]. Deletion is therefore one remedy among several; buyers should expect it to arrive with conduct bans, compliance reporting and data-retention limits.

Two caveats matter as of October 2026. These are settled consent orders, not court rulings, so they show what the Commission has negotiated rather than a judicially defined limit. The staff posts cited here also reflect the 2024 Commission's priorities, and current enforcement emphasis may differ [7]. For the underlying authority, see our overview of FTC Act Section 5 and AI data licensing.

What triggers model deletion exposure

The common trigger is a gap between what people were told and what the data was used for. Section 5 cases turn on deception (a false or misleading promise) or unfairness (substantial injury not reasonably avoidable and not outweighed by benefits), and training on personal data is where those gaps now surface.

The FTC's Office of Technology warned that model-as-a-service companies can be liable if they break promises not to use customer data for undisclosed purposes such as training [7]. A follow-up post said that quietly rewriting terms of service or privacy policies to allow AI training on data already collected could itself be unfair or deceptive [8]. The pattern applies upstream: if your supplier's users were promised one thing and you train on their records for another, your model inherits the problem.

Typical failure modes buyers should screen for:

  • Purpose drift. Records collected "to provide the service" are repurposed for third-party model training without updated notice or consent.
  • Retroactive policy changes. The supplier added an AI-training clause in 2024 but is licensing records collected under the 2021 policy [8].
  • Sensitive categories. Face geometry, voiceprints, health or children's data collected without the specific consent those regimes demand.
  • Deactivated or deleted accounts. Data that should have been purged under the supplier's own retention promise but remained in exports [3].
  • Untested downstream use. A model deployed in a consequential setting without accuracy testing, which compounds the data problem with an unfairness theory [6].

Why lineage determines how much you lose

The scope of a deletion order, and of your own remediation, depends on whether you can prove which models a dataset reached. An order phrased as "models or algorithms developed in whole or in part" with the affected data can sweep in every descendant checkpoint if you cannot show otherwise [1][3]. Without lineage, the safe reading of "in part" becomes "everything trained after the data landed."

A training-data register links each dataset version to the jobs, checkpoints and deployed endpoints that consumed it. Record content hashes of the shards, the job ID, the base checkpoint, and every fine-tune or distillation that inherited weights. Our audit readiness evidence pack covers the broader records regulators and customers request; the register below is the subset that scopes a disgorgement event.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy it matters in a deletion event
dataset_id / versionds-support-tickets v3.2Identifies the exact delivery under review
shard_hashesSHA-256 manifest of 412 Parquet filesProves which records were and were not present
source_notice_refSupplier privacy notice dated 2023-03-01, archived PDFTies collection to the promise people saw
consent_basisContract plus notice; no sensitive categoriesShows the theory you relied on
license_refLicense v1, permitted use: model trainingShows the supplier's grant and representations
consumed_by_jobsft-0917, ft-1002Starts the descendant tree
derived_checkpointsbase-7b-ft-0917, distill-1b-1010Everything that may need deletion or retraining
deployed_endpointssupport-assist-prodWhat has to come down or be swapped
deletion_statusNot triggeredRecords the outcome and date if invoked

Acquisition controls that reduce the risk

The cheapest disgorgement is the one you never face, and most of the work happens before data arrives. Check consents and notices at the source, get the supplier on the hook in writing, and refuse records whose origin cannot be explained.

Buyer checklist before accepting a training dataset with personal information:

  1. Read the actual notice. Obtain the privacy notice and terms in force when each record was collected, not just today's version, and check whether training or third-party sharing is disclosed. Our guide to consent and notice records lists what to request.
  2. Map collection dates to policy versions. Flag any record collected before an AI-use clause was added [8].
  3. Screen sensitive categories. Biometrics, health, financial and children's data each carry their own consent or de-identification standard.
  4. Get representations. Require the supplier to represent that it had the right to collect and license the data for training, and to notify you of any regulatory inquiry.
  5. Review the vendor's privacy program. Use a structured privacy review of the training data vendor covering retention, deletion and complaint handling.
  6. Negotiate deletion mechanics. Define what happens to derived models if the supplier's rights fail; see deletion and return clauses for licensed training data.
  7. Register before training. No dataset enters a training job until it has a register entry and a content-hash manifest. If you would rather source data whose rights and consents were reviewed before delivery, you can describe your data requirements to SourceX.

Removing direct identifiers before training helps but does not cure a collection problem. A deceptive-collection theory attaches to how data was obtained, and the photo-app order reached derived face embeddings and models, not only the raw images [3]. De-identified derivatives of improperly collected data can therefore still be in scope.

Planning remediation if an order or claim lands

Remediation cost scales with how entangled the data is, so plan the response before you need it. Expect three questions: which artifacts must go, how fast you can retrain, and whether anything short of retraining is credible.

Retraining cost. Removing data from a trained network usually means retraining from scratch, because a gradient-trained model has no record-level undo [9]. Keep the training configuration, data mixture and random seeds so a clean retrain on the remaining data is reproducible.

Checkpoint lineage. If the problem data entered only at a late fine-tuning stage, an earlier clean checkpoint may let you redo that stage alone. This is why the register should record the base checkpoint for every job.

Unlearning limits. Approximate unlearning methods exist, and SISA training (sharded, isolated, sliced, aggregated) reduces retraining cost by limiting each record's influence to one shard [9]. SISA must be adopted before training, and post-hoc approximate methods offer no exact deletion guarantee. Do not assume a regulator will accept an unlearning claim in place of the deletion an order specifies.

Retention conflicts. A deletion order can collide with documentation duties under other regimes. Our page on training data retention requirements covers how to keep the evidence while deleting the content.

How this fits wider disclosure duties

Disgorgement risk and transparency duties draw on the same records. California's AB 2013 requires developers of generative AI systems to post documentation about training data (disclosures were due by 1 January 2026), and the dataset-level facts you collect for that disclosure are the same facts that prove a model's lineage; see our AB 2013 training data disclosure guide. The compliance hub maps the remaining US and EU obligations that rely on a clean supplier record.

Sourcing data that reduces disgorgement exposure

SourceX sources operational datasets from US companies on request, reviews every dataset for ownership and consents, and delivers it under a license that defines the records, uses, term and delivery. Personal details such as names, emails and account numbers are removed or replaced before delivery and a sample is checked, though no method is perfect. Describe the data you need on the SourceX buyer page.

Sources

  1. Federal Trade Commission, "California Company Settles FTC Allegations It Deceived Consumers about use of Facial Recognition in Photo Storage App" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/01/california-company-settles-ftc-allegations-it-deceived-consumers-about-use-facial-recognition-photo
  2. Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
  3. Federal Trade Commission, "Analysis of Proposed Consent Order to Aid Public Comment (Everalbum)" (2021). https://www.ftc.gov/system/files/documents/cases/everalbum_analysis.pdf
  4. Federal Trade Commission, "Statement of Commissioner Rohit Chopra on Everalbum" (2021). https://www.ftc.gov/system/files/documents/public_statements/1585858/updated_final_chopra_statement_on_everalbum_for_circulation.pdf
  5. Federal Trade Commission, "The Great Doing: Remarks from the Chief Technologist Stephanie T. Nguyen (FAccT, June 2024)" (2024). https://www.ftc.gov/system/files/ftc_gov/pdf/stephanie-nguyen-remarks-facct-june-2024_0.pdf
  6. Federal Trade Commission (Business Guidance blog), "Coming face to face with Rite Aid's allegedly unfair use of facial recognition technology" (2023). https://www.ftc.gov/business-guidance/blog/2023/12/coming-face-face-rite-aids-allegedly-unfair-use-facial-recognition-technology
  7. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  8. Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  9. Bourtoule et al. (arXiv; IEEE Symposium on Security and Privacy), "Machine Unlearning" (2019). https://arxiv.org/abs/1912.03817v2

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data