Schemas, packaging and delivery
Propagating Deletions and Corrections Through Recurring Deliveries
Quick answer
Handle deletions in an incremental data feed by making them first-class records: the supplier ships a withdrawal list (stable record IDs, a reason code, an effective date, and whether the action is delete or correct), your pipeline applies it to the raw landing copy, then walks lineage to every derived shard, filtered subset, embedding index and cache. Keep a purge receipt per notice. Whether models already trained on withdrawn records must be retrained is a contract question, not a pipeline default.
By SourceX Editorial · Updated
This page covers mid-term removal and correction. For the mechanics of shipping only what changed, see incremental deliveries vs full refreshes; for end-of-term disposal, see certificates of data destruction. Both sit in the schemas, packaging and delivery hub.
Why withdrawals after delivery happen in licensed operational data
Withdrawals happen because the supplier's own source of truth keeps moving after your copy was cut. In operational data such as support tickets, CRM histories or finance workflows, the common triggers are a customer exercising a deletion right, a supplier's client asking to be excluded, a record found to be misattributed or under-redacted, and a contract or consent scope that turns out narrower than first thought.
Privacy law makes some of these non-negotiable. Under the CCPA, a business that receives a verifiable deletion request must also notify its service providers and contractors, and certain third parties, to delete the consumer's personal information [7]. In the EU, the EDPB has taken the position that a model trained on personal data is not automatically anonymous, and that unlawful processing at the development stage can affect how the model may later be used [8]. That makes downstream propagation an engineering requirement, not a courtesy.
Corrections are the quieter sibling. A ticket re-labeled from "billing" to "fraud," a repaired timestamp, or a redaction miss replaced with a placeholder all mean that rows already in your training mix are now wrong. Owner guidance on what happens when licensed data contains errors and on excluding specific clients from licensed data covers the supplier-side view.
How the withdrawal notice should be structured
A withdrawal notice should be a machine-readable file delivered through the same channel as the data, keyed by the same stable record IDs, never a prose email. If the supplier cannot point to an ID that survives across deliveries, nothing below works; the stable record IDs guide explains how to require one.
Change-data-capture systems already model this. Delta Lake's change data feed labels each row with a _change_type of insert, update_preimage, update_postimage or delete [1], and Debezium change events carry an operation field distinguishing create, update, delete and snapshot reads [2]. A licensed-data feed can borrow the same vocabulary even if it ships as files rather than a stream.
Illustrative example: invented to show structure; it does not describe an available dataset.
{"notice_id": "WN-2026-10-003", "dataset_version": "v2026.09", "record_id": "tkt_8f3a91c2", "action": "delete", "reason_code": "SUBJECT_ERASURE", "effective_date": "2026-10-01", "scope": "record_and_children", "child_tables": ["ticket_messages", "ticket_attachments"], "replacement_record_id": null}
{"notice_id": "WN-2026-10-003", "dataset_version": "v2026.09", "record_id": "tkt_5b07de11", "action": "correct", "reason_code": "REDACTION_MISS", "effective_date": "2026-10-01", "scope": "fields", "fields": ["body_text"], "replacement_record_id": "tkt_5b07de11@v2"}
Shipping this as JSON Lines keeps it appendable and diff-friendly, since each line is one UTF-8 JSON value [4]. The fields that matter most are listed below.
| Field | Purpose | Failure mode if missing |
|---|---|---|
record_id | Stable join key into every copy | Fuzzy matching on text, which misses near-duplicates |
action | delete, correct, or restrict (keep, but exclude from training) | Corrections applied as deletes, shrinking the corpus silently |
reason_code | Erasure, client exclusion, quality defect, scope change | Cannot decide whether models need review |
effective_date | When the record stops being usable | Disputes over whether a training run was compliant |
scope | Record only, record plus child rows, or specific fields | Orphaned messages or attachments remain after parent deletion |
notice_id | Ties every action to one auditable instruction | No way to prove which notice a purge answered |
Agree the reason-code list and notice cadence in the data contract for recurring deliveries, alongside schema rules.
Applying deletions to the files you actually store
Applying a deletion means rewriting the physical files that contain the record, because most training formats are not built for in-place removal. A Parquet file writes its metadata in a footer after the data, recording where each column chunk starts [3]; dropping a row means writing a new file and swapping the reference. A table format such as Delta Lake handles this with a transaction log and a DELETE command, but the old data files remain until vacuumed, so a logical delete is not a physical purge.
Sharded training formats are harder. WebDataset packs samples into tar shards, often around 1 GB each [5], so removing one conversation means repacking a shard and updating any shard index or sample-count manifest. JSONL is the simplest case: filter lines by record_id and rewrite, then regenerate checksums so the delivery manifest still verifies.
Watch three common traps:
- Time travel and snapshots. Iceberg snapshots, Delta history, S3 object versions and backup tiers keep the withdrawn bytes reachable until expired.
- Child records. Deleting a ticket header while leaving its messages, attachments or extracted entities behind leaves the content in place.
- Corrections as new versions. Store the corrected record under a new version ID and retire the old one, consistent with how you version licensed datasets.
Using lineage to reach derived copies
Lineage is what turns one withdrawal into a complete purge, because the raw landing zone is rarely where most copies live. A single support ticket can end up in a deduplicated corpus, a language-filtered subset, a tokenized and packed training shard, an eval holdout, an embedding in a vector index, a cached feature table, and a few notebooks.
Emit lineage events from every job that reads licensed data. OpenLineage run events record each job run with its input and output datasets [6], which lets you query "every dataset derived from licensed.vendor_x.tickets since v2026.07" instead of searching by hand. For record-level precision, carry the source record_id (or a list of contributing IDs for merged samples) as a column through each transform; dataset-level lineage alone tells you which files to check, not which rows to drop.
The same trace extends to models. If you already trace which models trained on which licensed dataset, you can answer the next question quickly: which checkpoints saw this record before its effective date.
Deciding whether trained models need action
Whether a model must be retrained, fine-tuned away from, or left alone after a withdrawal is decided by the license and applicable law, not by the pipeline. Some agreements require only that withdrawn records leave stored data and future training runs; others require action on models trained after a notice date; privacy-driven erasures may push further, given the EDPB's view on model anonymity [8].
The engineering choices determine how expensive each answer is. Full retraining is the reference method. SISA training, proposed by Bourtoule et al., shards and slices training data so that removing a point requires retraining only the affected constituent models, at some cost to accuracy [9]. For most production LLM work, the practical lever is cadence: if you retrain or refresh adapters on a schedule, a withdrawal can be honored at the next scheduled run, provided the contract allows that window.
Settle this before signing. The subscription data license refresh guide covers how refresh obligations are typically written, and what cancellation means for a data license covers the end-of-relationship case.
Proving each purge for audits
Each withdrawal notice should close with a purge receipt that a reviewer can check without trusting the engineer who ran it. Auditors, counsel and the supplier will ask the same three things: what was removed, from where, and when.
Illustrative example: invented to show structure; it does not describe an available dataset.
Purge receipt checklist
- Notice ID and receipt timestamp; record count by action (delete, correct, restrict)
- Raw landing copy rewritten; old object versions and table snapshots expired
- Every lineage descendant listed, with rewritten file paths and new checksums
- Vector index entries removed by
record_idmetadata filter, with before and after counts - Caches, feature stores and scratch buckets cleared or confirmed out of scope
- Model checkpoints that saw the records identified, with the contractually required action noted
- Post-purge query showing zero matches for the withdrawn IDs across all registered datasets
- Receipt signed off and retained with the delivery's diligence file
Restrict who can recreate a purged copy. The access controls for licensed training data guide covers bucket and role design that keeps ad-hoc exports from escaping lineage in the first place.
What to ask a supplier before the first delivery
Ask how withdrawals will reach you before you accept a recurring feed, because retrofitting a notice format after three deliveries means reconciling IDs by hand. Useful questions:
- Will deletions and corrections arrive as a separate notice file, as tombstone rows in the incremental delivery, or both?
- Are record IDs stable across deliveries and across parent and child tables?
- What reason codes exist, and does any of them imply model-level obligations?
- Is there a notice cadence, and how is an urgent erasure flagged?
- Can a correction be expressed at field level, or only as a full record replacement?
If you source recurring data through a partner, settle these questions in the license. SourceX sources operational datasets from US companies on request, rights-reviews each dataset for ownership and consents, and delivers under a license that defines records, uses, term and delivery; nothing is contracted until a supplier agrees. You can describe your recurring-data and withdrawal-handling needs through SourceX's buyer intake.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Scoping a recurring dataset with withdrawal handling
SourceX manages the commercial process for ongoing purchases, from assessing data and licensing permissions through agreeing pricing and allowed uses in a license. Every release is approved by the supplying company, and delivery runs through private, access-controlled workflows after an executed agreement. Describe the data you need and how you plan to keep it current at sourcex.si/buyers.
Sources
- Delta Lake (Linux Foundation), "Change data feed". https://docs.delta.io/delta-change-data-feed/
- Debezium project, "Debezium event changes". https://debezium.io/documentation/reference/transformations/event-changes.html
- The Apache Software Foundation (Apache Parquet), "File Format". https://parquet.apache.org/docs/file-format/
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- Hugging Face, "WebDataset (Hub documentation)". https://huggingface.co/docs/hub/datasets-webdataset
- OpenLineage (LF AI & Data), "OpenLineage spec examples". https://openlineage.io/docs/spec/examples
- California Privacy Protection Agency, "California Consumer Privacy Act of 2018 (statute text)". https://cppa.ca.gov/regulations/pdf/ccpa_statute.pdf
- European Data Protection Board, "Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models" (2024). https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
- Bourtoule et al., IEEE Symposium on Security and Privacy (arXiv), "Machine Unlearning" (2019). https://arxiv.org/abs/1912.03817v2
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.