Skip to content

Schemas, packaging and delivery

Iceberg and Delta Lake Tables as a Delivery Format for Licensed Data

Quick answer

Both Apache Iceberg and Delta Lake work well for licensed ML training data, because both keep data files (usually Parquet) under versioned table metadata that lets a training run pin an exact table version. Iceberg tracks columns by ID and has engine-neutral catalogs. Delta has the simpler path on Spark and Databricks. For licensed data, the format matters less than how you handle withdrawals. A row-level delete hides a row, but the bytes stay in storage until you expire snapshots or run VACUUM. Plan that step before the first delivery lands.

By SourceX Editorial · Updated

Why a table format beats loose files for licensed data

A table format gives you three things that loose Parquet or CSV files do not: atomic versions, a schema you can enforce, and row-level corrections. Each of these maps to an obligation that comes with a license. A supplier's first full load plus later corrections become a sequence of commits instead of a folder of part-0000*.parquet files with unclear precedence. Both formats usually store data as Parquet files underneath (Iceberg also supports ORC and Avro), so Parquet tools can inspect the files [5], although reading files directly ignores deletes and version boundaries.

The trade-off is operational. A delivered Iceberg or Delta table is only as trustworthy as its metadata, so you need a catalog, a maintenance schedule and someone who owns retention settings. Many teams therefore receive files under a manifest, verify them, and then commit them into their own table. Our guide on verifying a delivery with manifests and checksums covers that step. The delivery formats hub compares the alternatives.

Iceberg vs Delta Lake for ML training data: the working comparison

For most ML buyers, the choice follows the engines you already run, not the features. Iceberg favors multi-engine estates (Spark, Trino, Flink, Snowflake, BigQuery, Athena) that share a REST or Glue catalog. Delta favors Spark-centric and Databricks estates. The table below compares the behaviors that matter for licensed data. The table reflects the projects' documentation as of October 2026; confirm each item against the release you actually run, because both projects publish versioned docs and change defaults.

ConcernApache IcebergDelta LakeWhat to check on a licensed table
Version pinningSnapshot IDs and timestamps; named tags and branches [1]Integer table versions and timestamps in _delta_logRecord the snapshot ID or version in the training run config
Schema identityColumn IDs, so rename and drop are metadata changes [2]Name-based by default; rename and drop need column mapping modeWhether the supplier renames columns between deliveries
Row-level changesCopy-on-write, or merge-on-read via position or equality deletes (v2) and deletion vectors (format v3)Copy-on-write, or deletion vectors where enabledWhich delete mode is on, since it decides where the bytes live
Change historyIncremental reads between snapshotsChange data feed with _change_type [3]Whether corrections arrive as diffs you can audit
Physical removalExpire snapshots, then remove orphan filesVACUUM after the retention window [7]Who runs it, how often, and what evidence it leaves
Sharing protocolREST catalog with vended credentialsDelta Sharing [4]Whether the supplier shares live or ships a copy

Pinning a training run to an exact snapshot

Pin every training and eval job to an immutable table reference, never to "latest". In Iceberg, each commit creates a snapshot, and the snapshot log underpins both reader isolation and time travel [1]. A tag such as licensed_v2026_10 gives that snapshot a human-readable name and can carry its own retention, so routine expiry does not remove it [1]. In Delta, you pin with VERSION AS OF or TIMESTAMP AS OF and store the version number with the run.

The failure mode is quiet. A pinned reference only works while the files behind it still exist. Iceberg snapshot expiry and Delta VACUUM both delete the data files behind old versions. After a vacuum, Delta time travel to versions older than the retention period stops working. If your eval set is a pinned snapshot, its retention has to be set on purpose, not left to the cluster default.

Write the reference into your lineage records next to the license identifier. A useful minimum is catalog, namespace, table, snapshot ID or version, commit timestamp and row count. These fields let you answer "which licensed records trained this checkpoint" without guessing. They also support the training-data integrity practices that NIST SP 800-218A adds to secure development for AI models [6].

How schema evolution behaves across supplier deliveries

Iceberg handles supplier-side schema drift more safely because it tracks every column by a permanent integer ID, not by name or position [2]. Renaming ticket_body to body_text is a metadata change. Old files still resolve correctly, and a dropped column's ID is never reused, so a new column with the old name cannot inherit stale data [2]. Type promotion is limited to safe widenings such as int to long and float to double.

Delta's default is name-based. Additive changes work through mergeSchema on write, and overwriteSchema replaces the schema outright. Renaming or dropping a column requires enabling column mapping (delta.columnMapping.mode = 'name'), which raises the table protocol version, and older readers may then fail to open the table. Check reader compatibility before you turn it on for a table other teams consume.

Whatever the format, do not let a supplier's delivery evolve your production schema automatically. Land each delivery in a staging table. Diff its schema against the agreed data dictionary, then promote it. Our delivery specification template has fields for declaring which changes are allowed between deliveries.

Row-level deletes for corrections and withdrawals

Row-level deletes are how you apply a supplier's correction or a record withdrawal, but a logical delete only hides the row. It does not remove it. This distinction matters most for licensed data, because a license may require you to stop using records or destroy them.

Iceberg format v2 offers two delete styles. Position deletes name a data file path and row position. Equality deletes name column values, such as source_record_id = 'R-88213', so the writer does not need to know where the row lives. Format v3 adds deletion vectors as per-file bitmaps stored in Puffin files [9]. With all three, the deleted bytes stay in storage until compaction rewrites the data file. Even after that, older snapshots still point to the original file.

Delta behaves the same way. A DELETE under copy-on-write rewrites the affected files, but the old files stay referenced by earlier versions. With deletion vectors enabled, the row is masked and the file is untouched. Either way, physical removal needs a later VACUUM.

Equality deletes keyed on a stable source identifier are the cleanest way to apply supplier withdrawals. That only works if the delivery has one. See stable record IDs across deliveries for how to request them.

From logical delete to physical removal: vacuum and snapshot expiry

Physical removal is a separate maintenance job, and it is what a destruction record should point to. In Delta, VACUUM deletes data files that the current table no longer references and that are older than the retention threshold. The default threshold is 7 days [7]. VACUUM does not run automatically, and it leaves _delta_log files in place. The docs warn against short retention intervals, because concurrent readers or streams can fail if their files disappear. Managed platforms add options such as dry runs. Treat those as platform features, not open-source guarantees.

In Iceberg, the sequence is to rewrite data files so that deleted rows are gone from current files, expire the snapshots that still reference the old files, and then remove orphan files. The Spark procedures are rewrite_data_files, expire_snapshots and remove_orphan_files. Tags and branches with longer retention keep their snapshots alive, so a forgotten tag on a withdrawn record defeats the whole run [1].

Neither format covers copies outside the table. Staging buckets, feature stores, cached shards, notebooks and object-store versioning or soft delete can all keep the bytes. A destruction step has to cover them as well.

Illustrative example: invented to show structure; it does not describe an available dataset.

Withdrawal-to-destruction runbook (lakehouse)

  1. Receive the withdrawal list keyed on source_record_id, with the license reference and the effective date.
  2. Apply the delete: an Iceberg equality delete or MERGE, or a Delta DELETE. Record the resulting snapshot ID or version.
  3. List every tag, branch and pinned eval snapshot that still contains the records, and decide whether to retire or rebuild each one.
  4. Iceberg: run rewrite_data_files on the affected partitions, then expire_snapshots older than the delete commit, then remove_orphan_files. Delta: if deletion vectors are on, run REORG TABLE ... APPLY (PURGE) [8], then VACUUM once the retention window has passed.
  5. Purge object-store noncurrent versions and any derived copies (shards, embeddings, feature tables).
  6. Log the commands, timestamps, before and after file listings, and the operator as evidence for the license file. Agree what destruction evidence the license expects before you sign; if you are scoping a purchase now, you can describe a data request to SourceX.

Catalogs, sharing and Delta–Iceberg interoperability

Choose the catalog first. It decides who can read the table and whether you receive a live share or a copy. Iceberg tables need a catalog to resolve the current metadata pointer, such as a REST catalog, AWS Glue, Hive Metastore, Nessie or Polaris. Delta keeps its log beside the data and can be shared through Delta Sharing, an open protocol documented in the Delta Lake project [4].

A live share leaves the supplier in control of versions and retention. That is convenient, but you cannot pin what the supplier vacuums. If you need reproducible training runs, materialize a copy into your own account. Cross-account bucket delivery and access controls after delivery cover the surrounding controls.

Interoperability layers such as Delta UniForm and Apache XTable generate Iceberg metadata for Delta tables, or translate between the formats, over the same Parquet files. They are useful for reading across engines. Feature support can lag the native format, however, so test deletes, column mapping and time travel through the translated view before relying on them. For ongoing purchases, also decide whether corrections arrive as table commits or as file drops. Incremental deliveries vs full refreshes covers that choice.

What to put in the delivery spec when you want a table

If you want a supplier to deliver a table, specify the table contract, not just the format name. Ask for:

  • Format and version (Iceberg format v2 or v3; Delta protocol reader and writer versions)
  • Partition spec and sort order
  • A primary key or stable record ID
  • The delete mode
  • Whether change data feed or incremental snapshots are enabled
  • The expected schema with column IDs or a column mapping mode
  • Retention settings on delivered tags
  • A manifest of data files with checksums

Many suppliers' source systems are relational databases or SaaS exports, so a Parquet drop that you commit yourself is often the practical path. For record-level detail on operational tables, see tabular and transactional data for AI and the data lakehouse glossary entry.

Sourcing tabular operational data for your lakehouse

SourceX sources operational datasets from US companies, such as support and sales histories, engineering records, and finance and legal workflows. It does so on request; it does not hold them in stock. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, through private, access-controlled workflows. Delivery happens only after an executed agreement and supplier approval. To describe the tables your training or eval pipeline needs, start at SourceX for AI data buyers.

Sources

  1. Apache Software Foundation (Apache Iceberg), "Branching and Tagging (Iceberg 1.7.0 documentation)". https://iceberg.apache.org/docs/1.7.0/branching
  2. Apache Software Foundation (Apache Iceberg), "Schema evolution (Iceberg docs, evolution.md)". https://apache.googlesource.com/iceberg/+show/refs/heads/1.1.x/docs/evolution.md
  3. Delta Lake project (Linux Foundation), "Change data feed". https://docs.delta.io/delta-change-data-feed/
  4. Delta Lake project (Linux Foundation), "Read Delta Sharing Tables". https://docs.delta.io/delta-sharing/
  5. Apache Software Foundation (Apache Parquet), "Apache Parquet Documentation". https://parquet.apache.org/docs
  6. National Institute of Standards and Technology, "Secure Software Development Practices for Generative AI and Dual-Use Foundation Models (NIST SP 800-218A)" (2024). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-218A.pdf
  7. Delta Lake, "Table utility commands (VACUUM)". https://docs.delta.io/3.0.0/delta-utility.html
  8. Microsoft, "Reorganize Delta tables with REORG (PURGE)". https://docs.databricks.com/aws/delta/vacuum
  9. Amazon Web Services, "Unlock the power of Apache Iceberg v3 deletion vectors on Amazon EMR". https://datalakehousehub.com/blog/iceberg-delete-story-2026

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data