Skip to content

Schemas, packaging and delivery

Handling Schema Changes Across Recurring Dataset Deliveries

Quick answer

Handle schema changes in recurring data deliveries by treating every delivery as a versioned schema, not a folder of files. Classify each change on intake as additive, breaking or semantic; gate the load with an automated diff against the last accepted version; and map columns by stable identifiers rather than by name or position. Contract the notice period and the change format with the supplier, and keep code-list drift (new enum values) in scope, because it corrupts training labels without breaking a single parser.

By SourceX Editorial · Updated

Which schema changes break a recurring feed?

A change breaks a feed when an existing reader can no longer interpret old or new records correctly, and that includes changes no parser will complain about. Engineers usually sort changes into three groups. Additive changes add a nullable column or optional field. Structural breaking changes rename, drop, retype or re-nest fields. Semantic changes keep the column but alter what its values mean.

The third group causes the most damage in training data. A ticket_status column that gains pending_vendor passes every type check, yet a classifier trained on the old five-value label set now sees an unmapped class. A resolution_minutes field that silently switches to seconds passes range checks if your bounds are loose. A free-text notes field that starts including agent macros changes the text distribution of a fine-tuning set.

For the delivery formats themselves, see the delivery cluster hub and the trade-offs in Parquet vs JSONL for licensed training data.

How schema registry compatibility modes map to file deliveries

Compatibility modes from schema registries give file-based feeds a precise vocabulary for what the supplier may change without notice. Confluent Schema Registry's documentation defines BACKWARD (readers on the new schema can read data written with the latest registered schema), FORWARD (readers on the latest registered schema can read data written with the new one), FULL (both), a _TRANSITIVE variant of each that checks against all prior versions rather than only the last, and NONE, which disables checks.

Confluent's compatibility table also states what each mode permits. BACKWARD allows deleting fields and adding optional fields and expects consumers to upgrade first; FORWARD allows adding fields and deleting optional fields and expects producers to upgrade first; FULL allows adding and deleting optional fields in any order. In Avro and Protobuf, a field with a default value can be added or removed as a fully compatible change, while JSON Schema has no compatibility rules of its own, so behavior depends on the registry's implementation. Check the current documentation for the registry you use before writing these rules into a contract [6].

For a buyer receiving monthly batches, the useful translation is this. You re-read history when you retrain, so you need BACKWARD_TRANSITIVE at minimum: today's reader must parse every delivery you still hold. If you also keep older pipelines running on frozen schemas, ask for FULL_TRANSITIVE.

Why column identity matters more than column names

Pipelines that join deliveries by column name or position are the main reason a rename turns into data loss. Apache Iceberg tracks each column by a unique ID, and its documentation notes that name-tracked formats can accidentally "un-delete" a column when a name is reused, while position-tracked formats struggle to drop a column without shifting the rest [7]. CSV and headerless TSV deliveries are position-tracked by default, which makes a dropped middle column the classic silent-corruption bug.

Parquet files carry a self-describing schema in their footer, and the normative format specification lives in the separate parquet-format repository [3]. That helps, but a supplier's export tool may still renumber or rename fields between runs. Ask the supplier to include a field ID or stable field key in the data dictionary, and keep a mapping table on your side so that cust_tier in March and customer_tier in April resolve to the same internal column.

The same logic applies to records: changes in keys are covered in stable record IDs and join keys across deliveries.

Detecting schema drift at intake

Detect drift by diffing each delivery's observed schema and value domains against the last accepted version before anything lands in a training table. Ingestion engines expose this as a policy choice. Microsoft's Unity Catalog training material contrasts a fail-on-new-columns mode, which halts the stream so a person can review the change, with evolve modes that add the column automatically, and covers notification when a change is detected [1]. AWS documents that its MSK data delivery to Iceberg does not evolve the table in place and points users to a new destination when the schema must change [2].

For licensed training data, default to fail-closed for structural changes and quarantine for semantic ones. Auto-evolve is acceptable only for nullable additive columns that no downstream feature or prompt template reads yet.

A practical intake diff covers five layers:

  1. Structure: field set, order (for CSV), nesting, nullability, physical and logical types (for example Parquet INT64 with a TIMESTAMP(MILLIS) annotation versus MICROS).
  2. Domains: distinct values for every enum-like field, compared with the last accepted code list.
  3. Units and encodings: timezone of timestamps, currency fields, text encoding, decimal precision and scale.
  4. Shape variance: fields whose type depends on context. Zendesk ticket audit Change events, for example, carry field_name, value and previous_value, and value is an array when the field is tags [5]; a flattening step that assumes a scalar will break on the first tag change.
  5. Distribution: null rates, string-length percentiles and category frequencies against a rolling baseline, which catches renamed-in-place semantics that pass type checks.

Encode the structural layer as a machine-readable contract. As of October 2026, JSON Schema 2020-12 is the current release, split into Core and Validation specifications [4], and works well for JSONL; for Parquet, compare the Arrow schema extracted from the footer. Deeper validation patterns live in validation checks for structured dataset deliveries and using JSON Schema to validate deliveries.

Decision table: classifying and handling each change type

Use a fixed decision table so that the same change gets the same response every delivery, regardless of who is on call.

Illustrative example: invented to show structure; it does not describe an available dataset.

Change observedClassRegistry analogueIntake actionTraining-data risk
New nullable column channel_typeAdditiveBACKWARD-safeAuto-accept; leave unmapped until reviewedLow, unless a template iterates over all columns
Column cust_tier renamed customer_tierBreaking (structural)Not compatible without aliasBlock; map via field-key table; reloadOld/new rows split into two features
amount changed from DECIMAL(12,2) to DOUBLEBreaking (type)Not compatibleBlock; cast explicitly with tolerance checkRounding drift in numeric targets
Column region droppedBreaking (removal)FORWARD-safe only if optionalBlock if any consumer reads itSilent nulls in features
status gains value pending_vendorSemantic (code list)Invisible to registriesQuarantine; extend label map; version taxonomyUnmapped or mislabeled class
Timestamps switch from UTC to local timeSemantic (units)Invisible to registriesBlock; require supplier confirmationOrdering and recency errors
notes begins including templated macrosSemantic (content)Invisible to registriesFlag via distribution check; decide filterText distribution shift in fine-tuning set

The last three rows are the reason schema validation alone is not enough. Registries reason about field presence and types; they say nothing about what a value means.

What the supplier's change notice should contain

Notice expectations belong in the contract, and the technical response is a versioned schema with a migration attached. The commercial side of that agreement (notice periods, remedies, acceptance) is covered in data contracts for recurring deliveries and ongoing data supply agreements. On the engineering side, ask that every change arrive as a structured notice alongside the delivery manifest.

Illustrative example: invented to show structure; it does not describe an available dataset.

schema_change_notice:
  dataset: support_tickets_monthly
  from_schema_version: "3.2.0"
  to_schema_version: "4.0.0"        # major bump = breaking
  first_delivery_with_change: "2026-11"
  changes:
    - type: rename
      field_key: f_017
      old_name: cust_tier
      new_name: customer_tier
    - type: code_list_extend
      field_key: f_009
      field: status
      added_values: [pending_vendor]
      meaning: "Waiting on a third party; previously coded as 'open'"
    - type: type_change
      field_key: f_022
      field: amount
      old_type: "DECIMAL(12,2)"
      new_type: "DECIMAL(14,4)"
  backfill: "Prior 12 months re-exported under 4.0.0 in same delivery"
  data_dictionary_ref: dictionary_v4.0.0.csv

Two fields matter most. meaning on a code-list change tells you whether historical records need relabeling (here, some old open tickets were really pending_vendor). backfill tells you whether you can rebuild a consistent history or must carry a version boundary into training. Keep the notice with the snapshot; see versioning licensed datasets for release-note practice.

Absorbing the change without corrupting training and eval sets

Absorb breaking changes by migrating into a stable internal schema rather than letting the supplier's schema propagate into training tables. Keep three layers: a raw landing zone stored exactly as delivered and tagged with the supplier schema version, a conformed layer where versioned migrations map every supplier version into your canonical columns, and the training or eval views built from the conformed layer only.

Write each migration as code keyed by (from_version, to_version) and test it against a frozen sample of the prior delivery. If the supplier cannot backfill, add a schema_version column to the conformed layer and decide explicitly whether to train across the boundary. For eval sets, freeze the label taxonomy per benchmark version; a new status value should create a new eval version, not silently alter scores on the old one.

Corrections and deletions that arrive alongside a schema change need their own handling; see propagating deletions and corrections. For refresh cadence, SourceX's notes on how often AI buyers want fresh data and the scenario on delivering fresh data each quarter frame how often these checks will run.

Sourcing recurring feeds with change handling in mind

When you describe the data you need to SourceX, include your expected schema, refresh cadence and change-handling requirements alongside the use case. SourceX sources operational datasets from US companies on request, manages the commercial process, including licensing agreements and ongoing purchases, and delivers every rights-reviewed, supplier-approved release under a license that defines records, uses, term and delivery. Describe the recurring feed you need.

Frequently asked questions

Is adding a column always safe?

No. It is safe for readers that select columns explicitly, but SELECT jobs, CSV readers bound to positions and prompt templates that serialize every field will change behavior. Auto-accept only nullable columns and keep them unmapped until someone reviews them.

How do we handle a field renamed in a data feed if the supplier gave no notice?

Block the load, confirm the rename with the supplier and record the mapping in your field-key table. Matching on a near-identical name plus a matching value distribution is a reasonable heuristic for triage, but treat it as a hypothesis until confirmed.

Does Parquet handle schema evolution for us?

Partly. Parquet stores the schema with each file [3], so readers can merge files with different column sets, but merge behavior for renames and type changes depends on the engine. Table formats such as Iceberg add ID-based column tracking on top.

Sources

  1. Microsoft Learn, "Detect and manage schema drift". https://learn.microsoft.com/en-us/training/modules/implement-manage-data-quality-constraints-unity-catalog/4-detect-manage-schema-drift
  2. Amazon Web Services, "Troubleshoot schema changes in MSK data delivery". https://docs.aws.amazon.com/msk/latest/developerguide/msk-data-delivery-ts-schema-change.md
  3. The Apache Software Foundation (Apache Parquet), "Apache Parquet Documentation". https://parquet.apache.org/docs
  4. JSON Schema, "JSON Schema Specification". https://json-schema.org/specification
  5. Zendesk Developer Docs, "Ticket Audit events reference". https://developer.zendesk.com/documentation/ticketing/reference-guides/ticket-audit-events-reference/
  6. Confluent, "Schema Evolution and Compatibility". https://docs.confluent.io/platform/current/schema-registry/fundamentals/schema-evolution.html
  7. Apache Software Foundation, "Schema Evolution - Apache Iceberg". https://iceberg.apache.org/docs/latest/evolution/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data