Schemas, packaging and delivery
Certificates of Data Destruction for Licensed Datasets
Quick answer
A certificate of data destruction for a licensed dataset is a signed attestation that every copy of named dataset versions was deleted or returned, by a stated method, on a stated date, with any exceptions listed. To be credible it must trace back to an inventory of copies: buckets, warehouse tables, derived shards, caches, laptops and backups. It should cite a recognized method such as NIST SP 800-88 [1], and say plainly what the license requires for trained models.
By SourceX Editorial · Updated
What a destruction certificate for licensed data has to prove
The certificate has to prove scope, method and accountability, not just that "the data was deleted." IT asset disposal certificates were built around one physical drive with a serial number, and vendors describe them as auditable records of a specific process applied to a specific asset at a specific time [3]. A licensed dataset is different because it is a logical asset that has been copied, transformed and cached across many systems since delivery.
That shifts the burden to inventory. A licensor reading your certificate will ask three questions. Which copies existed, how was each one removed, and who signed with knowledge of the facts? University data-destruction attestations show the institutional pattern: a named custodian formally certifies that sensitive data was fully destroyed under stated requirements [2].
The obligation itself comes from the license, so read the deletion and termination clauses before you start. Our guide to AI data license terms covering use, exclusivity and deletion explains the clauses; this page covers the evidence. For the policy side, see retention and deletion in data licensing.
Building the copy inventory before anything is deleted
Start from the delivery manifest and walk forward through every system that touched the files. If you verified the delivery against a dataset manifest with checksums, those SHA-256 hashes are your best tool for finding copies, because a renamed file still hashes the same. Version IDs from your dataset versioning scheme tell you which releases are in scope.
Copies hide in predictable places. Check these locations, which are the ones most often missed in close-out reviews:
- Landing and staging: the original receiving bucket, SFTP drop directories, and any decrypted working copy left after PGP or age decryption.
- Warehouse and lakehouse: loaded tables plus their time-travel and fail-safe retention, cloned tables, and shares created through zero-copy warehouse sharing.
- Derived artifacts: tokenized shards, deduplicated subsets, embedding indexes in a vector store, train/eval splits, and filtered JSONL or Parquet exports.
- Compute-side caches: Hugging Face
datasetscache directories, notebook scratch volumes, container images that baked in sample files, and dataloader caches on training nodes. - Endpoints and tickets: laptops used for exploration, samples pasted into issue trackers or chat, and spreadsheets attached to evaluation reviews.
- Backups and replicas: snapshot schedules, cross-region replication targets, and backup vaults with their own retention.
The access logs you kept under post-delivery access controls help here. Every principal that read the data is a lead for a copy you have not found yet.
Choosing Clear, Purge or Destroy under NIST SP 800-88
NIST SP 800-88 gives the vocabulary most licensors and auditors expect: Clear, Purge and Destroy. As of October 2026, Rev. 2 is the final revision and supersedes Rev. 1 from 2014; check the NIST CSRC record for the current text [1]. Clear uses logical techniques such as overwriting user-addressable storage, Purge resists laboratory-level recovery (for example cryptographic erase or block erase), and Destroy renders media physically unusable [1].
Vendor summaries of Rev. 2 highlight its coverage of flash and cloud storage and its fuller treatment of cryptographic erase, including key sanitization [1]. That matters for licensed data, because you rarely control the physical media in a cloud account. In practice, a dataset sitting in a cloud bucket encrypted under a customer-managed KMS key you own can be cryptographically erased by deleting the objects and then scheduling destruction of a dedicated key. That only works if the key was used for that dataset alone, which is a good reason to assign a per-dataset key at intake.
Deleting objects or tables in a cloud account is logical deletion, not Clear in the SP 800-88 sense, because you cannot overwrite the provider's media. Record it as logical deletion and rely on the provider's own media sanitization for the hardware.
Match the method to the location:
| Location | Practical method | SP 800-88 category | Evidence to keep |
|---|---|---|---|
| Object storage, versioned bucket | Delete every version ID and delete marker; expire noncurrent versions | Logical deletion (no SP 800-88 category) | Version listing before and after, lifecycle rule JSON |
| Object storage under dedicated CMK | Object deletion plus scheduled key destruction | Purge (cryptographic erase) | KMS key ID, deletion schedule, CloudTrail or audit log event |
| Warehouse tables | DROP plus expiry of time-travel and fail-safe windows | Logical deletion (no SP 800-88 category) | Query history, retention settings, date windows close |
| Laptop SSD | MDM erase that destroys the device encryption key, or full device wipe | Purge (cryptographic erase) | MDM wipe record, device serial |
| Retired drives or encrypted shipping drives | Physical shredding by a disposal vendor | Destroy | Vendor certificate with serial numbers |
Cloud deletion semantics that break naive certificates
A "deleted" object in cloud storage is often still there, and certificates fail when the signer does not know this. In an Amazon S3 bucket with versioning enabled, a DELETE request without a version ID only places a delete marker on top of the object; earlier versions remain and can be restored by removing the marker. Data is permanently removed only when each specific versionId is deleted.
Lifecycle rules behave the same way. The Expiration action on current versions adds a delete marker and keeps the data as a noncurrent version, while NoncurrentVersionExpiration is the action that permanently removes noncurrent versions. If MFA Delete is enabled, version-level deletes need the x-amz-mfa header, so plan who holds that device. Other providers have equivalents in soft delete, object retention locks and warehouse time travel; record the retention window for each and certify only after it has closed.
Two more traps are common. S3 replication does not replicate deletions of specific object versions [6], so you must purge each replica bucket separately. Object Lock in compliance mode blocks deletion of a version until its retain-until date, and an Object Lock legal hold blocks it until the hold is removed; both belong in the exceptions section rather than being glossed over.
Backups, legal holds and trained weights
Backups are the hardest copies to certify, and the honest approach is to certify the expiry schedule rather than pretend an immutable backup was edited. Most backup vaults cannot delete one dataset from inside a snapshot. State the backup system, the snapshot IDs that contain the data, the retention policy, and the date the last affected snapshot expires, and confirm that restores during that window are blocked for the licensed data.
Legal holds override deletion. If litigation or a regulator's request requires preservation, list the held copy, its custodian and the hold reference, and commit to a follow-up certificate when the hold lifts. Statutes can also impose their own destruction duties: the Illinois Biometric Information Privacy Act requires entities holding biometric data to publish a retention schedule and guidelines for permanent destruction [5].
Trained weights are where many licenses are silent. Commentary on AI contract termination notes that clauses commonly require return or deletion of the data but often do not say whether models already trained on it may continue in use [4]. Your certificate should restate whatever the license says about models, checkpoints and embedding indexes, and list model and checkpoint IDs in scope. Dataset-to-model traceability records make that list defensible. For the broader question, see whether AI labs delete data after training.
A certificate of data destruction template for licensed datasets
A usable certificate is short on prose and long on identifiers. The template below shows the fields a licensor's counsel or auditor will look for; adapt field names to your contract's defined terms.
Illustrative example: invented to show structure; it does not describe an available dataset.
certificate_id: CDD-2026-0412
license_reference: "Data License Agreement dated 2025-03-01, Section 9.3"
licensee: "Example Labs Inc. (data operations)"
dataset:
name: "support-ticket-histories"
versions_in_scope: ["v2025.03.0", "v2025.09.1"]
manifest_sha256: "e3b0c442...b855" # hash of each delivered manifest
disposition: destroy # destroy | return | return_then_destroy
copies:
- location: "s3://example-landing/support-tickets/"
method: "delete all versionIds + delete markers; NoncurrentVersionExpiration=1"
sp800_88: none # logical deletion
completed: 2026-09-30
evidence: ["version-listing-before.json", "version-listing-after.json"]
- location: "warehouse: analytics.support_tickets_v2"
method: "DROP TABLE; time-travel window closed 2026-10-07"
sp800_88: none # logical deletion
completed: 2026-10-07
- location: "kms key alias/ds-support-tickets"
method: "key scheduled for deletion; waiting period ended 2026-10-08"
sp800_88: purge
completed: 2026-10-08
- location: "laptop serial C02XXXX (MDM)"
method: "remote wipe"
sp800_88: purge
completed: 2026-09-29
exceptions:
- item: "nightly backup snapshots containing v2025.09.1"
reason: "immutable backups"
expires: 2026-11-15
restore_blocked: true
- item: "none under legal hold"
models_and_derivatives:
checkpoints_in_scope: ["ckpt-sft-0817", "ckpt-sft-0902"]
treatment: "as required by license Section 9.4" # restate the clause
attestation:
signer: "Name, Head of Data Operations"
basis: "personal review of inventory and evidence listed above"
date: 2026-10-09
followup_certificate_due: 2026-11-16
The structure does three things a generic disposal certificate does not. It names dataset versions and manifest hashes, so the licensor can match it to what was delivered. It separates completed deletions from dated exceptions, and it commits to a follow-up when the last backup expires.
Return versus destroy, and how to evidence each
Return and destroy produce different evidence, and some licenses require both in sequence. A return is evidenced like a delivery in reverse: transfer logs, a checksum match on the receiving side, and the licensor's acknowledgment. Only after that acknowledgment should you start destruction, otherwise you may delete your only proof of what was returned.
Keep the evidence pack separate from the data. Version listings, KMS audit events, query history exports, MDM wipe records and disposal vendor certificates should sit in a records system with its own retention policy, so you can answer an audit request years later. Do not include samples of the licensed data in the evidence pack; hashes and identifiers are enough.
When the license close-out is part of an ongoing purchase rather than an end of relationship, the same process applies per version. Retiring an old snapshot under an incremental or full-refresh delivery model is a smaller certificate with the same fields. The delivery hub for licensed AI data links the other intake and packaging guides that feed this inventory.
Where SourceX fits in license close-out
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, and delivery runs through private, access-controlled workflows after an executed agreement and supplier approval. That gives your close-out team defined records and a defined delivery to certify against. Buyers can describe the data they need on the SourceX buyers page.
For the wider licensing context, see the buyer's guide to AI training data licensing and the guide to data provenance for AI training data.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Get licensed data you can certify at the end of the term
SourceX looks for US businesses that hold the data you describe, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that sets out records, uses, term and delivery. Tell SourceX what data your team needs.
Sources
- NIST, "SP 800-88 Rev. 2, Guidelines for Media Sanitization". https://csrc.nist.gov/pubs/sp/800/88/r2/final
- Brown University IT, "Certificate of Data Destruction". https://ithelp.brown.edu/kb/articles/pdf/certificate-of-data-destruction
- SECURIS, "What is a Certificate of Data Destruction?". https://securis.com/blog/what-is-a-certificate-of-data-destruction/
- Reed Smith LLP, "Contractual considerations: security, performance and termination (Entertainment and Media Guide to AI)". https://www.reedsmith.com/articles/entertainment-and-media-guide-to-ai/contractual-considerations-security-performance-and-termination/
- Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
- Amazon Web Services, "Object Lock considerations". https://docs.aws.amazon.com/AmazonS3/latest/userguide/object-lock-managing.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.