Skip to content

Schemas, packaging and delivery

Versioning Licensed Datasets: Snapshots, Version IDs and Release Notes

Quick answer

Dataset versioning best practice for licensed data comes down to four commitments. Every delivery becomes an immutable, content-addressed snapshot. Each snapshot gets a version ID that encodes what kind of change it carries. Each version ships with a release note listing added, removed and corrected records and any schema change. Training runs and license records both reference that exact ID. Done this way, any model can name the data it used, and any license can name the deliveries it covers.

By SourceX Editorial · Updated

Why licensed data needs stricter versioning than internal data

Licensed data needs stricter versioning because the version ID becomes a legal reference as well as an engineering one. With internal data, a lost version costs reproducibility. With licensed data, it can also leave you unable to show which records a term covered, which deliveries a deletion notice applies to, or whether a model trained before or after a supplier's correction.

The Research Data Alliance's versioning recommendations, drawn from 39 use cases, center on citing the exact version used and attributing each version to its source [1]. A conceptual framework from ANU separates six concerns that buyers tend to collapse into one "version" field: Revision, Release, Granularity, Manifestation, Provenance and Citation [2]. For a buyer, the useful reading is that a supplier "release" (what they sent) and your "manifestation" (how you stored it, say Parquet after converting from CSV) are different objects and need linked but separate IDs.

Vendor guidance from ML tooling firms frames the same goal more practically: versions exist so you can rerun an experiment on identical data and trace a performance shift back to a dataset change [4].

Immutable snapshots: the unit everything else points to

An immutable snapshot is a read-only, checksummed copy of a delivery that is never edited in place; corrections arrive as a new snapshot. This is the rule that makes every downstream reference trustworthy. If someone can overwrite s3://licensed/acme-tickets/v1.2.0/ after a run, the version ID in your experiment tracker means nothing.

Three implementation patterns are common:

  • Object storage plus manifest. Land each delivery under a version-scoped prefix, write a manifest of file paths, byte sizes and SHA-256 hashes, and enable object lock or a deny-delete bucket policy. Pair this with manifest and checksum verification at intake.
  • Table formats. Apache Iceberg keeps a snapshot log that supports time travel, and tags give a named, fixed reference to one snapshot with its own retention setting. Tagging acme_tickets_v1_2_0 against the snapshot ID created by the load is a clean way to make a table-format version citable.
  • Data version control tools. DVC and lakeFS add commit-style history over files. The choice between them is covered in DVC vs lakeFS vs table formats; this page stays with the scheme, not the tool.

One failure mode recurs across all three: snapshot expiry. Iceberg's expire_snapshots removes unused snapshots and data files according to retention properties, so a default housekeeping job can silently delete the snapshot a production model was trained on. Set retention on tagged releases explicitly and exclude them from routine cleanup unless a license or deletion obligation says otherwise [6].

Choosing a version ID scheme: semantic, date-based or hybrid

Use semantic versioning when deliveries change shape, date-based IDs when they only change content, and a hybrid when both happen. Semantic Versioning 2.0.0 defines MAJOR.MINOR.PATCH and requires the project to define a public API before the numbers carry meaning. For a dataset, that "public API" is the schema and the documented semantics in the data dictionary: field names, types, enumerations, ID stability and the meaning of labels.

A workable mapping for licensed datasets:

  • MAJOR: a breaking change to consumers. A dropped or renamed field, a changed type, a redefined label taxonomy, a new record ID scheme, or a re-processing pass (new redaction method, new tokenization) that changes text content across the board.
  • MINOR: additive and backward compatible. New records from a refresh window, a new nullable column, a new enum value appended to an open list.
  • PATCH: corrections that do not change shape. Fixed mislabels, removed duplicates, records withdrawn after a supplier deletion request.

SemVer also allows pre-release labels such as -rc.1 and build metadata after a plus sign, and build metadata does not affect version precedence [5]. Use pre-release tags for supplier samples and acceptance candidates (2.0.0-rc.1) and build metadata for your own manifestation (2.0.0+parquet.zstd), so conversion never looks like a new supplier release.

Date-based IDs (2026-09-30 or 2026.09) suit append-only feeds where the schema is pinned by a data contract and every delivery is a window. The weakness is that a date says nothing about compatibility, so teams end up reading release notes to learn whether a pipeline will break. A hybrid such as 2.3.0 with a window_end: 2026-09-30 field in the manifest keeps both signals.

What a dataset release note must contain

A release note is the human- and machine-readable record of what changed between two versions, and it should be complete enough to decide, without opening the data, whether a training run must be redone. Ask suppliers for it as a delivery requirement in your technical delivery specification, and generate your own diff to check it.

Minimum contents:

  • Record deltas. Counts and ID lists of added, removed and corrected records, keyed on stable record IDs. Without stable IDs, "corrected" is indistinguishable from "removed plus added".
  • Reason codes for removals. Supplier withdrawal, deletion request, duplicate, out of scope, failed de-identification check. Removals tied to deletion requests must flow into your deletion and correction propagation process.
  • Schema changes. Added, renamed, retyped or deprecated fields, with a link to the schema evolution policy that classified them.
  • Re-processing notes. Any change to de-identification method, label guidelines, transcription model or normalization, since these alter content without altering IDs.
  • Coverage. Source time window, systems of record, and known gaps.
  • Integrity. Manifest hash for the whole release and the parent version ID.

If you publish an internal catalog entry per version, a Hugging Face-style dataset card works well as the template: the README carries narrative notes while a YAML block at the top records structured metadata such as license, language and size [3]. Add version, parent_version and license_ref keys to that block.

Illustrative example: invented to show structure; it does not describe an available dataset.

dataset: support-tickets-saas-b2b
version: 2.1.0
parent_version: 2.0.1
release_date: 2026-09-30
window: { start: 2026-07-01, end: 2026-09-30 }
manifest_sha256: 9f2c...e41a
license_ref: LIC-0142 / Schedule B, Delivery 3
changes:
  added_records: 18412
  removed_records: { count: 37, reasons: { deletion_request: 21, duplicate: 16 } }
  corrected_records: { count: 512, fields: [resolution_code] }
  schema: [ "added nullable field: escalation_tier (string, enum open)" ]
  reprocessing: none
compatibility: backward-compatible (MINOR)
superseded_versions_retained: [2.0.1]

Pinning a version to every training and eval run

Every training and evaluation run should record the dataset version ID and manifest hash it read, not a path or a "latest" alias. Aliases such as latest or a mutable Iceberg branch are convenient for exploration but resolve differently over time, which defeats reproducibility.

Practical rules for an ML platform team:

  1. Data loaders accept only tagged versions or snapshot IDs in production configs; an unpinned reference fails the job.
  2. The experiment tracker logs dataset_id, version, manifest_sha256 and any filter or split seed, because a filter applied at load time is effectively a new manifestation.
  3. Eval sets get their own version line, separate from training data, so contamination checks can compare record IDs across both.
  4. Model registry entries inherit the list of dataset versions from their runs. This is the backbone of dataset-to-model traceability.

The RDA's emphasis on citing the exact version used [1] maps directly onto this: a model card or internal audit entry that says "trained on support tickets" is not a citation, while "support-tickets-saas-b2b 2.1.0, manifest 9f2c...e41a" is.

Tying license scope to versions

License scope should be expressed in terms of deliveries and versions, and the mapping between license schedules and version IDs should live in your data catalog, not only in a contract PDF. Licenses typically define the records covered, permitted uses, term and delivery mechanics. When refresh deliveries arrive, each new version needs a recorded answer to "which license and which schedule covers this?"

A simple decision table helps engineering and counsel stay aligned:

Illustrative example: invented to show structure; it does not describe an available dataset.

Change in new versionVersion bumpLicense question to checkCatalog action
New records from next refresh windowMINORIs this window inside the term and the delivery schedule?Link version to schedule entry
Records withdrawn on a deletion requestPATCHMust prior versions and derived artifacts also be purged?Flag superseded versions for remediation
New field with new data type (e.g., call audio)MAJORDoes the license's record definition include it?Hold from training until confirmed
Re-processed with a new de-identification methodMAJORDoes the method still meet the agreed standard?Attach the method record to the version
Format conversion on your sideBuild metadata onlyNone, unless the license limits copiesRecord manifestation, keep parent link

For recurring purchases, this mapping is what makes ongoing data supply agreements auditable. It also gives you the evidence trail described in SourceX's guide to keeping a record of what you licensed. How often new versions arrive is a separate design question; see how often AI buyers want fresh data and the trade-offs in incremental deliveries vs full refreshes.

Retaining superseded versions versus deletion obligations

Keep superseded versions only as long as reproducibility needs and the license allow, and treat deletion obligations as overriding snapshot retention. Immutability and deletion are in tension: the snapshot that guarantees reproducibility is also a copy that a term end, a supplier withdrawal or a data subject request may require you to destroy.

Resolve it with explicit policy rather than ad hoc exceptions:

  • Record a retention class per version (active, superseded-retained, superseded-pending-deletion, destroyed) in the catalog.
  • When a PATCH release removes records for a deletion request, scan older versions and derived artifacts (tokenized shards, embeddings, eval subsets) for the same record IDs.
  • When a term ends, expire tags and run snapshot expiry deliberately, then verify the underlying objects are gone, including backups and replicas. Document the outcome as covered in certificates of data destruction.
  • Keep the metadata (manifest hashes, release notes, license mapping) after deleting the data, unless the license says otherwise; it is what lets you prove what you held and when.

Restrict who can read retained superseded versions using the same controls as active data; see access controls for licensed training data.

How SourceX fits versioned, recurring deliveries

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Each dataset is rights-reviewed and delivered under a license that defines the records, uses, term and delivery, through private, access-controlled workflows after an executed agreement and supplier approval. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset. Buyers can describe the data they need on the SourceX buyers page; a request does not guarantee a match. More delivery guidance sits in the schemas, packaging and delivery hub and the wider AI data guide library.

Request licensed data you can version from day one

SourceX finds US businesses that hold the operational data you describe and manages the path from assessment through a license that defines records, uses, term and delivery. Nothing is contracted until a supplier agrees, and every release is approved by the supplying company. Describe the licensed dataset you need.

Sources

  1. Research Data Alliance, "Principles and Best Practices in Data Versioning for All Data Sets Big and Small". https://archive.rd-alliance.org/sites/default/files/Principles%20and%20Best%20Practices%20in%20Data%20Versioning.pdf
  2. Australian National University Open Research, "Versioning data is about more than revisions: A conceptual framework and proposed principles". https://openresearch-repository.anu.edu.au/items/790eff8e-907e-4dd8-b259-85f4c18a9144
  3. Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
  4. Latitude, "Best Practices for Dataset Version Control". https://latitude.so/blog/best-practices-for-dataset-version-control
  5. Tom Preston-Werner, "Semantic Versioning 2.0.0". https://semver.org/
  6. Apache Software Foundation, "Maintenance - Apache Iceberg". https://iceberg.apache.org/docs/latest/maintenance/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data