Skip to content

Schemas, packaging and delivery

Certificates of Data Destruction for Licensed Datasets

Quick answer

A certificate of data destruction for a licensed dataset is a signed attestation that every copy of named dataset versions was deleted or returned, by a stated method, on a stated date, with any exceptions listed. To be credible it must trace back to an inventory of copies: buckets, warehouse tables, derived shards, caches, laptops and backups. It should cite a recognized method such as NIST SP 800-88 [1], and say plainly what the license requires for trained models.

By SourceX Editorial · Updated

What a destruction certificate for licensed data has to prove

The certificate has to prove scope, method and accountability, not just that "the data was deleted." IT asset disposal certificates were built around one physical drive with a serial number, and vendors describe them as auditable records of a specific process applied to a specific asset at a specific time [3]. A licensed dataset is different because it is a logical asset that has been copied, transformed and cached across many systems since delivery.

That shifts the burden to inventory. A licensor reading your certificate will ask three questions. Which copies existed, how was each one removed, and who signed with knowledge of the facts? University data-destruction attestations show the institutional pattern: a named custodian formally certifies that sensitive data was fully destroyed under stated requirements [2].

The obligation itself comes from the license, so read the deletion and termination clauses before you start. Our guide to AI data license terms covering use, exclusivity and deletion explains the clauses; this page covers the evidence. For the policy side, see retention and deletion in data licensing.

Building the copy inventory before anything is deleted

Start from the delivery manifest and walk forward through every system that touched the files. If you verified the delivery against a dataset manifest with checksums, those SHA-256 hashes are your best tool for finding copies, because a renamed file still hashes the same. Version IDs from your dataset versioning scheme tell you which releases are in scope.

Copies hide in predictable places. Check these locations, which are the ones most often missed in close-out reviews:

  • Landing and staging: the original receiving bucket, SFTP drop directories, and any decrypted working copy left after PGP or age decryption.
  • Warehouse and lakehouse: loaded tables plus their time-travel and fail-safe retention, cloned tables, and shares created through zero-copy warehouse sharing.
  • Derived artifacts: tokenized shards, deduplicated subsets, embedding indexes in a vector store, train/eval splits, and filtered JSONL or Parquet exports.
  • Compute-side caches: Hugging Face datasets cache directories, notebook scratch volumes, container images that baked in sample files, and dataloader caches on training nodes.
  • Endpoints and tickets: laptops used for exploration, samples pasted into issue trackers or chat, and spreadsheets attached to evaluation reviews.
  • Backups and replicas: snapshot schedules, cross-region replication targets, and backup vaults with their own retention.

The access logs you kept under post-delivery access controls help here. Every principal that read the data is a lead for a copy you have not found yet.

Choosing Clear, Purge or Destroy under NIST SP 800-88

NIST SP 800-88 gives the vocabulary most licensors and auditors expect: Clear, Purge and Destroy. As of October 2026, Rev. 2 is the final revision and supersedes Rev. 1 from 2014; check the NIST CSRC record for the current text [1]. Clear uses logical techniques such as overwriting user-addressable storage, Purge resists laboratory-level recovery (for example cryptographic erase or block erase), and Destroy renders media physically unusable [1].

Vendor summaries of Rev. 2 highlight its coverage of flash and cloud storage and its fuller treatment of cryptographic erase, including key sanitization [1]. That matters for licensed data, because you rarely control the physical media in a cloud account. In practice, a dataset sitting in a cloud bucket encrypted under a customer-managed KMS key you own can be cryptographically erased by deleting the objects and then scheduling destruction of a dedicated key. That only works if the key was used for that dataset alone, which is a good reason to assign a per-dataset key at intake.

Deleting objects or tables in a cloud account is logical deletion, not Clear in the SP 800-88 sense, because you cannot overwrite the provider's media. Record it as logical deletion and rely on the provider's own media sanitization for the hardware.

Match the method to the location:

LocationPractical methodSP 800-88 categoryEvidence to keep
Object storage, versioned bucketDelete every version ID and delete marker; expire noncurrent versionsLogical deletion (no SP 800-88 category)Version listing before and after, lifecycle rule JSON
Object storage under dedicated CMKObject deletion plus scheduled key destructionPurge (cryptographic erase)KMS key ID, deletion schedule, CloudTrail or audit log event
Warehouse tablesDROP plus expiry of time-travel and fail-safe windowsLogical deletion (no SP 800-88 category)Query history, retention settings, date windows close
Laptop SSDMDM erase that destroys the device encryption key, or full device wipePurge (cryptographic erase)MDM wipe record, device serial
Retired drives or encrypted shipping drivesPhysical shredding by a disposal vendorDestroyVendor certificate with serial numbers

Cloud deletion semantics that break naive certificates

A "deleted" object in cloud storage is often still there, and certificates fail when the signer does not know this. In an Amazon S3 bucket with versioning enabled, a DELETE request without a version ID only places a delete marker on top of the object; earlier versions remain and can be restored by removing the marker. Data is permanently removed only when each specific versionId is deleted.

Lifecycle rules behave the same way. The Expiration action on current versions adds a delete marker and keeps the data as a noncurrent version, while NoncurrentVersionExpiration is the action that permanently removes noncurrent versions. If MFA Delete is enabled, version-level deletes need the x-amz-mfa header, so plan who holds that device. Other providers have equivalents in soft delete, object retention locks and warehouse time travel; record the retention window for each and certify only after it has closed.

Two more traps are common. S3 replication does not replicate deletions of specific object versions [6], so you must purge each replica bucket separately. Object Lock in compliance mode blocks deletion of a version until its retain-until date, and an Object Lock legal hold blocks it until the hold is removed; both belong in the exceptions section rather than being glossed over.

Backups are the hardest copies to certify, and the honest approach is to certify the expiry schedule rather than pretend an immutable backup was edited. Most backup vaults cannot delete one dataset from inside a snapshot. State the backup system, the snapshot IDs that contain the data, the retention policy, and the date the last affected snapshot expires, and confirm that restores during that window are blocked for the licensed data.

Legal holds override deletion. If litigation or a regulator's request requires preservation, list the held copy, its custodian and the hold reference, and commit to a follow-up certificate when the hold lifts. Statutes can also impose their own destruction duties: the Illinois Biometric Information Privacy Act requires entities holding biometric data to publish a retention schedule and guidelines for permanent destruction [5].

Trained weights are where many licenses are silent. Commentary on AI contract termination notes that clauses commonly require return or deletion of the data but often do not say whether models already trained on it may continue in use [4]. Your certificate should restate whatever the license says about models, checkpoints and embedding indexes, and list model and checkpoint IDs in scope. Dataset-to-model traceability records make that list defensible. For the broader question, see whether AI labs delete data after training.

A certificate of data destruction template for licensed datasets

A usable certificate is short on prose and long on identifiers. The template below shows the fields a licensor's counsel or auditor will look for; adapt field names to your contract's defined terms.

Illustrative example: invented to show structure; it does not describe an available dataset.

certificate_id: CDD-2026-0412
license_reference: "Data License Agreement dated 2025-03-01, Section 9.3"
licensee: "Example Labs Inc. (data operations)"
dataset:
  name: "support-ticket-histories"
  versions_in_scope: ["v2025.03.0", "v2025.09.1"]
  manifest_sha256: "e3b0c442...b855"   # hash of each delivered manifest
disposition: destroy              # destroy | return | return_then_destroy
copies:
  - location: "s3://example-landing/support-tickets/"
    method: "delete all versionIds + delete markers; NoncurrentVersionExpiration=1"
    sp800_88: none   # logical deletion
    completed: 2026-09-30
    evidence: ["version-listing-before.json", "version-listing-after.json"]
  - location: "warehouse: analytics.support_tickets_v2"
    method: "DROP TABLE; time-travel window closed 2026-10-07"
    sp800_88: none   # logical deletion
    completed: 2026-10-07
  - location: "kms key alias/ds-support-tickets"
    method: "key scheduled for deletion; waiting period ended 2026-10-08"
    sp800_88: purge
    completed: 2026-10-08
  - location: "laptop serial C02XXXX (MDM)"
    method: "remote wipe"
    sp800_88: purge
    completed: 2026-09-29
exceptions:
  - item: "nightly backup snapshots containing v2025.09.1"
    reason: "immutable backups"
    expires: 2026-11-15
    restore_blocked: true
  - item: "none under legal hold"
models_and_derivatives:
  checkpoints_in_scope: ["ckpt-sft-0817", "ckpt-sft-0902"]
  treatment: "as required by license Section 9.4"   # restate the clause
attestation:
  signer: "Name, Head of Data Operations"
  basis: "personal review of inventory and evidence listed above"
  date: 2026-10-09
followup_certificate_due: 2026-11-16

The structure does three things a generic disposal certificate does not. It names dataset versions and manifest hashes, so the licensor can match it to what was delivered. It separates completed deletions from dated exceptions, and it commits to a follow-up when the last backup expires.

Return versus destroy, and how to evidence each

Return and destroy produce different evidence, and some licenses require both in sequence. A return is evidenced like a delivery in reverse: transfer logs, a checksum match on the receiving side, and the licensor's acknowledgment. Only after that acknowledgment should you start destruction, otherwise you may delete your only proof of what was returned.

Keep the evidence pack separate from the data. Version listings, KMS audit events, query history exports, MDM wipe records and disposal vendor certificates should sit in a records system with its own retention policy, so you can answer an audit request years later. Do not include samples of the licensed data in the evidence pack; hashes and identifiers are enough.

When the license close-out is part of an ongoing purchase rather than an end of relationship, the same process applies per version. Retiring an old snapshot under an incremental or full-refresh delivery model is a smaller certificate with the same fields. The delivery hub for licensed AI data links the other intake and packaging guides that feed this inventory.

Where SourceX fits in license close-out

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, and delivery runs through private, access-controlled workflows after an executed agreement and supplier approval. That gives your close-out team defined records and a defined delivery to certify against. Buyers can describe the data they need on the SourceX buyers page.

For the wider licensing context, see the buyer's guide to AI training data licensing and the guide to data provenance for AI training data.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Get licensed data you can certify at the end of the term

SourceX looks for US businesses that hold the data you describe, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that sets out records, uses, term and delivery. Tell SourceX what data your team needs.

Sources

  1. NIST, "SP 800-88 Rev. 2, Guidelines for Media Sanitization". https://csrc.nist.gov/pubs/sp/800/88/r2/final
  2. Brown University IT, "Certificate of Data Destruction". https://ithelp.brown.edu/kb/articles/pdf/certificate-of-data-destruction
  3. SECURIS, "What is a Certificate of Data Destruction?". https://securis.com/blog/what-is-a-certificate-of-data-destruction/
  4. Reed Smith LLP, "Contractual considerations: security, performance and termination (Entertainment and Media Guide to AI)". https://www.reedsmith.com/articles/entertainment-and-media-guide-to-ai/contractual-considerations-security-performance-and-termination/
  5. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  6. Amazon Web Services, "Object Lock considerations". https://docs.aws.amazon.com/AmazonS3/latest/userguide/object-lock-managing.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data