Skip to content

Schemas, packaging and delivery

DVC vs lakeFS vs Table Formats: Versioning Tools for Received Datasets

Quick answer

Pick DVC when datasets are file collections tied to a git-managed training repo, lakeFS when many teams need git-style branches, commits and merges directly over an S3, Azure or GCS bucket, and Apache Iceberg or Delta Lake when the data is tabular and queried by Spark, Trino or a warehouse. For licensed data, the deciding question is not just reproducibility but deletion: each tool keeps history by design, and each needs a specific, tested purge procedure.

By Noah Loul, Founder & CEO, SourceX · Updated

This page compares the tools for one job: receiving a licensed dataset, pinning the exact version a model trained on, and removing records when a license or supplier requires it. For the broader practice of version IDs, release notes and snapshot naming, see versioning licensed datasets; this page owns the tool choice.

How each tool actually stores a version

Each tool stores history differently, and that storage model determines both scale limits and how hard a deletion is. Get this layer right before comparing features.

DVC keeps small metafiles (.dvc files and dvc.lock) in git that record content hashes, while the data lives in a content-addressed cache and one or more remotes such as S3, GCS, Azure Blob or SSH. A git commit therefore pins a dataset version indirectly, through hashes. Nothing about access control lives in DVC itself; it inherits whatever the remote bucket enforces.

lakeFS sits in front of object storage as a versioning layer with an S3-compatible gateway. Repositories have branches, commits, tags and merges, and objects are kept by default so you can return to earlier versions. Readers address data as s3://repo/branch-or-commit/path, so a training job can pin a commit ID without copying data.

Iceberg and Delta Lake version at the table level. Iceberg tables keep a snapshot log that supports reader isolation and time travel, and named branches and tags point at specific snapshots with their own retention settings [1]. Delta keeps a transaction log in _delta_log and can optionally write a change data feed with a _change_type column for inserts, updates and deletes [2]. Both fit naturally into a data lakehouse.

Git LFS replaces large files in a git repo with pointer files and stores content on an LFS server. It works for modest assets, but hosted LFS services impose storage and bandwidth quotas, every clone negotiates object downloads through the LFS server, and removing a file from history requires rewriting git history. For multi-terabyte deliveries it is usually the wrong tool.

Comparison table: DVC, lakeFS, Iceberg/Delta and Git LFS

The short version is that DVC is a project tool, lakeFS is a bucket-level tool, and table formats are a query-engine tool. The table compares them on the dimensions an ML platform team usually weighs.

DimensionDVClakeFSIceberg / Delta LakeGit LFS
Unit of versionFiles and directories by content hashObjects in a repository commitTable snapshotIndividual files
Where history livesgit metafiles + cache/remotelakeFS metadata + storage namespaceTable metadata / transaction loggit + LFS server
Best fitOne repo, file datasets, pipelines (dvc.yaml)Many teams sharing buckets, branch-per-experimentTabular data read by Spark, Trino, Flink, warehousesSmall binaries next to code
Pin a training run bygit commit SHAlakeFS commit ID or tagSnapshot ID, tag, or timestampgit commit SHA
Access controlBucket IAM on the remotelakeFS RBAC plus storage IAMCatalog and engine permissionsRepo permissions
Extra infrastructureNone beyond a remotelakeFS server and metadata storeCatalog (REST, Glue, Unity, Hive)LFS server or host quota
Hard-delete mechanismdvc gc with scope [3] and --cloudRetention rules + GC job [4]Expire snapshots / VACUUMHistory rewrite + server-side purge

Scale tends to follow the same order. DVC handles large datasets well when files are reasonably sized, but millions of tiny files make hashing and dvc status slow; pack them into shards first (see streaming formats in object storage). lakeFS commits are metadata operations, so branching a large repository does not copy data. Table formats scale to very large tables but only version what is in the table, not loose PDFs or audio.

Purging required deletions from immutable history

Every one of these tools retains old versions until you run an explicit collection step, so deleting a record from the current version does not delete it from storage. If a supplier withdraws records, a data subject request flows through, or a license ends, plan the purge for each layer separately.

DVC. Remove the files from the workspace, commit, then run dvc gc. The command [3] does nothing unless you give a scope such as --workspace, --all-branches, --all-tags or --all-commits, and --cloud is needed to delete from the remote as well as the local cache. The scope is the trap: if any retained branch, tag or commit still references the hash, --all-commits keeps it. Remote deletion is irreversible unless another remote or backup holds the data, and a shared cache can break links for other projects unless you pass --projects. Old git commits still contain the hashes, which is metadata, not content, but record that in your deletion log.

lakeFS. Delete the objects on every branch, commit, then let garbage collection hard-delete them. Retention [4] is set per branch with a repository default, and an object shared across branches is removed only when retention expires for all of them. Objects reachable from any branch HEAD are never collected, so a forgotten experiment branch silently blocks deletion. After collection, commits remain readable but a read of a purged object returns HTTP 410 Gone, which your training loaders should handle. Two caveats: objects brought in with lakectl import are not touched by GC, and, as of October 2026, hosted and enterprise tiers run managed or standalone GC instead of the Spark job.

Iceberg and Delta. Run a row-level DELETE, then remove the snapshots that still reference the old data files. In Iceberg, tags and branches carry their own retention, so a tag you created to pin a training run keeps the deleted rows alive until you drop or expire it [1]. In Delta, VACUUM removes unreferenced files older than the retention threshold, and if change data feed is enabled, the feed files keep a delete row image of what you removed until they are vacuumed too [2].

Watch for soft deletes: Iceberg merge-on-read tables write delete files and Delta tables with deletion vectors mark rows instead of rewriting files, so compact or rewrite the affected data files (for Delta, REORG TABLE ... APPLY (PURGE)) before expiring or vacuuming. Time travel to a pre-deletion version must stop working before you can call the deletion complete.

Copies outside the tool. Local DVC caches on GPU nodes, lakeFS-exported snapshots, warehouse clones and object-store versioning (S3 versioned buckets keep noncurrent versions) all survive the tool-level purge. Inventory them through your access controls for licensed training data and evidence the result in a certificate of data destruction.

Decision checklist for choosing a versioning tool

The right tool usually follows from three facts: data shape, number of consumers, and how often deletions are expected. Work through the checklist with the license terms beside you.

Illustrative example: invented to show structure; it does not describe an available dataset.

QuestionIf yesIf no
Is the delivery mainly tabular (Parquet, CSV) queried with SQL?Iceberg or Delta, pinned by snapshot tagContinue
Do several teams train from the same bucket and need isolated branches?lakeFSContinue
Is the data a file collection used by one training repo?DVC with a dedicated remote per licenseContinue
Are files under a few hundred MB total and rarely changed?Git LFS is acceptableAvoid Git LFS
Does the license require deletion within a fixed window?Write and test the purge runbook before first loadStill document GC settings
Will ongoing purchases add incremental drops?Prefer append-friendly table formats or lakeFS commits per dropSingle snapshot is enough

A common, workable combination is lakeFS or a dedicated bucket for raw deliveries, a table format for curated tabular layers, and DVC metafiles in the training repo to pin which version each experiment used. Keep one license per remote or repository where possible; mixing licenses in one DVC cache or one lakeFS repository makes scoped garbage collection far harder.

Pinning versions for reproducible training and eval

Reproducibility means a run record that names the immutable version, not a branch. Record the DVC commit SHA, lakeFS commit ID or Iceberg snapshot ID, plus the delivery manifest hash, in your experiment tracker.

Illustrative example: invented to show structure; it does not describe an available dataset.

run_id: ft-2026-10-07-ops-tickets-v3
dataset:
  license_ref: LIC-0042
  delivery_id: drop-2026-09
  manifest_sha256: 9f2c...e41a
  versioning:
    tool: lakefs
    repository: licensed-support-tickets
    commit_id: c7a91f0e
    tag: train-ft-2026-10
  deletion_events_applied: [del-2026-09-30]
eval_set:
  tool: iceberg
  table: eval.support_holdout
  snapshot_id: 5823417096412

Two failure modes recur. Tags created to pin runs are the most common blocker for deletion, because they keep snapshots and objects alive by design. And branch names like main are not pins; a run recorded only as "main" cannot be reproduced after the next commit. Verify bytes on arrival with a manifest and checksums and link each model to its inputs through dataset-to-model traceability.

Where the delivery format fits in

The versioning tool should match how data arrives, so agree the delivery shape before choosing. Columnar drops in Parquet rather than JSONL load cleanly into Iceberg or Delta, while document, audio and email deliveries are better as file collections under DVC or lakeFS.

Recurring deliveries add schema drift, which table formats handle through schema evolution and file-based tools do not; see schema changes across recurring deliveries and incremental vs full-refresh deliveries. The delivery hub covers the rest of the packaging and transfer stack, and the AI data hub covers sourcing more broadly.

SourceX sources operational datasets from US companies on request and manages licensing and ongoing purchases, with every dataset delivered under a license defining records, uses, term and delivery. If you are choosing tooling for an incoming dataset, you can describe the data you need to SourceX and plan the versioning around the license.

Getting licensed datasets ready to version

SourceX rights-reviews every dataset for ownership and consents and delivers it under a license that defines records, uses, term and delivery, through private, access-controlled workflows after an executed agreement and supplier approval. Data is sourced on request rather than held in stock, and a request does not guarantee a match. Tell SourceX what data your team needs.

Sources

  1. Apache Software Foundation (Apache Iceberg), "Branching and Tagging (Apache Iceberg 1.7.0)". https://iceberg.apache.org/docs/1.7.0/branching
  2. Delta Lake (Linux Foundation), "Change data feed". https://docs.delta.io/delta-change-data-feed/
  3. Data Version Control (DVC), "gc". https://doc.dvc.org/command-reference/gc
  4. lakeFS, "Garbage Collection". https://docs.lakefs.io/admin/garbage-collection/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data