Skip to content

Schemas, packaging and delivery

Data Contracts for Recurring Dataset Deliveries

Quick answer

A data contract specification for a recurring dataset delivery is a versioned, machine-readable file, usually YAML in the Open Data Contract Standard (ODCS), that states the feed's schema, field meanings, quality rules, delivery location and change policy. Both sides keep it in version control. The buyer runs it automatically against every drop, quarantining anything that fails before the data reaches a training, eval or RAG pipeline. Commercial remedies stay in the license; the contract is their technical counterpart.

By SourceX Editorial · Updated

What a data contract covers in a supplier feed

A data contract covers everything a pipeline needs to accept or reject a delivery without a human reading a PDF. The Open Data Contract Standard, an open YAML specification maintained by the Bitol project under the Linux Foundation, groups this into fundamentals, schema, data quality rules, servers, ownership, SLA properties, support channels and custom properties. Its v3 line publishes a JSON Schema against which the contract file itself can be validated. Verify section and field names against the current spec on the Bitol site before adopting them.

For a feed that arrives weekly or monthly, the sections that matter most are:

  • Fundamentals: contract id, version, status (draft, active, deprecated), the dataset name and a plain-language purpose.
  • Schema: each object (table or file family) with properties, logicalType, physicalType, required, primaryKey and classification.
  • Semantics: a description per field, units, time zone, enum meanings and the business definition behind derived fields. This is where a data dictionary moves into the contract (see the data dictionary template for dataset deliveries).
  • Quality: declarative rules such as null limits, uniqueness, value ranges and row-count bounds.
  • Servers: where each drop lands (bucket, prefix, warehouse share) and in which format.
  • Ownership and support: a named role on each side and the channel for contract change requests.

One nuance matters between companies. ODCS is usually written by the producer as a description of what a dataset should look like, not as a bilateral agreement. For a licensed feed, treat the file as a jointly reviewed annex: the supplier publishes it, the buyer accepts a specific version, and the license refers to that version.

Illustrative ODCS-style contract for a support-ticket feed

A working contract is short enough to read in one sitting and precise enough for a validator to enforce. The example below shows the shape for a monthly feed of de-identified support tickets used to refresh a RAG corpus and a held-out eval set. Field names follow ODCS v3 conventions; check them against the published JSON Schema for the version you pin before you adopt them. A dataset-level metadata format such as Croissant can sit alongside it to describe the corpus for ML tooling, but it does not replace per-delivery acceptance rules [5].

Illustrative example: invented to show structure; it does not describe an available dataset.

apiVersion: v3.1.0
kind: DataContract
id: urn:contract:acme-support-tickets
name: support_tickets_monthly
version: 2.1.0            # semver: MAJOR = breaking, MINOR = additive
status: active
description:
  purpose: Monthly de-identified support tickets for RAG and eval refresh
  usage: Governed by license schedule B; this file is technical only
servers:
  - server: buyer-landing
    type: s3
    location: s3://buyer-landing/supplier-x/tickets/
    format: parquet
schema:
  - name: tickets
    physicalName: tickets
    properties:
      - name: ticket_id
        logicalType: string
        required: true
        primaryKey: true
        description: Stable supplier key; never reused across deliveries
      - name: opened_at
        logicalType: date
        physicalType: timestamp
        required: true
        description: UTC, ISO 8601
      - name: channel
        logicalType: string
        description: email | chat | phone_transcript
      - name: body_redacted
        logicalType: string
        required: true
        classification: restricted
        description: Ticket text after PII replacement with typed placeholders
      - name: resolution_code
        logicalType: string
    quality:
      - type: library
        metric: rowCount
        mustBeGreaterThan: 1000
      - type: library
        metric: duplicateValues
        arguments: { properties: [ticket_id] }
        mustBe: 0
slaProperties:
  - property: frequency
    value: monthly
  - property: latency
    value: 5
    unit: d
team:
  - role: supplier_data_owner
  - role: buyer_data_ops

Three choices in this file prevent most recurring-feed incidents. ticket_id is declared as a stable primary key, which is what lets you process upserts and deletions (see stable record IDs across deliveries). opened_at names its time zone. body_redacted carries a classification so downstream access controls can key off it.

Writing a change policy both sides can enforce

A change policy works when it classifies every possible schema change in advance and ties each class to a version bump and a notice period. Semantic versioning is the usual vocabulary: a MAJOR bump for anything that can break a consumer, MINOR for additive changes, PATCH for description-only edits. Contract tooling that diffs two contract versions and sorts the differences by severity makes the version bump checkable rather than a matter of trust.

Illustrative example: invented to show structure; it does not describe an available dataset.

ChangeClassVersion bumpBuyer pipeline behavior
Add optional propertyAdditiveMINORAccept; new column ignored until mapped
Add required propertyBreakingMAJORQuarantine until contract version accepted
Remove or rename propertyBreakingMAJORQuarantine
Narrow a type (string to enum, wide to narrow numeric)BreakingMAJORQuarantine
Change an enum's meaning without changing its valuesBreaking (semantic)MAJORQuarantine; needs human review
Tighten a quality thresholdNon-breaking for buyerMINORAccept
Edit descriptions onlyEditorialPATCHAccept

Semantic changes are the dangerous row because no validator sees them. If resolution_code "R3" used to mean "refund issued" and now means "escalated," every label derived from it silently flips. Require that enum meanings live in the contract so the diff surfaces them.

Renames deserve explicit rules. Table formats such as Apache Iceberg track columns by permanent IDs so a rename does not orphan data [2], but CSV and JSONL drops have no such mechanism. For file feeds, treat a rename as remove plus add. The page on handling schema changes across recurring deliveries covers migration mechanics; the contract's job is to declare which class each change falls into.

Validating every delivery in CI

Contract validation runs in two stages: lint the contract file, then test the delivered data against it. The open-source Data Contract CLI (datacontract) is one common tool for ODCS files: a lint command checks the YAML against the standard and a test command runs schema and quality checks against the server the contract names, with credentials passed as environment variables. Command names have changed between releases, so confirm them against the version you install and pin that version in CI.

A practical gate for each drop:

  1. Pin the version. Read the contract version the supplier declares in the delivery manifest and refuse drops that reference a version you have not accepted.
  2. Verify bytes first. Check checksums and file counts from the dataset manifest before any schema test, so a truncated upload is not misreported as a quality failure.
  3. Lint, then test. Run datacontract lint on the pinned file, then datacontract test against the landing prefix.
  4. Diff on contract updates. When the supplier proposes a new version, diff the old and new files with your tooling's changelog or breaking-change command and fail the job on any MAJOR-class difference that arrives without a MAJOR bump.
  5. Quarantine, do not drop. Move failing deliveries to a quarantine prefix with the test report attached, and keep promotion to the training zone as a separate, logged step.
  6. Record lineage. Emit a run event for each validation so you can tie a model build back to an accepted delivery; OpenLineage models this as jobs, runs and datasets [4].

Streaming or auto-ingest paths need the same discipline at the reader. Databricks Auto Loader's failOnNewColumns mode, for example, stops a stream when a column outside the defined schema appears [3], which is the streaming equivalent of a MAJOR-change quarantine.

Where the contract ends and the license begins

The data contract defines what a valid delivery looks like; the license defines what happens when a delivery is not valid. Freshness targets, cure periods, credits and termination rights belong in the agreement and in your SLA metrics for recurring deliveries. ODCS can carry SLA properties such as frequency and latency, which is useful for measurement, but the remedy must still sit in the signed document. The data supplier SLAs guide covers the service terms themselves.

Keep three links explicit. The license schedule should name the contract id and accepted version. The contract description.usage should point back to that schedule instead of restating permitted uses, so the two never drift. And the change notice period for MAJOR versions should appear in the license, because a YAML file cannot bind anyone to wait.

Failure modes seen in recurring-feed contracts

Most data contract failures come from the contract being too loose, too stale or validated in the wrong place. The common ones:

  • Types only, no semantics. logicalType: string passes every value; without enum lists, units and time zones, the contract catches format drift but not meaning drift.
  • Contract edited in place. A supplier updates version: 2.1.0 without bumping it. Hash the accepted contract file and compare on every run.
  • Checks on the supplier side only. A supplier's passing test proves nothing about what landed in your bucket after transfer and decryption; run tests on your landing copy.
  • No deletion semantics. Incremental feeds need a declared tombstone or deletion file format; see propagating deletions and corrections.
  • Quality thresholds set once. A rowCount floor that fit month one may hide a 40% drop in month nine. Review thresholds with each MINOR version.
  • JSON Schema drafts mixed. If you also validate JSONL records with JSON Schema, pin a draft; as of October 2026, 2020-12 is the current release, split into Core and Validation [1]. The JSON Schema validation guide covers record-level checks.

Refresh cadence also shapes how strict you need to be: a daily feed can tolerate a quarantine-and-wait loop that a quarterly one cannot. The question of how often AI buyers want fresh data is worth settling before you set thresholds.

Sourcing recurring operational feeds with SourceX

If you are planning a recurring feed, it helps to arrive with a draft contract in hand, because it turns "we need support data" into a checkable specification. SourceX sources operational datasets, such as support and sales histories, engineering records and finance and legal workflows, from US companies on request, and manages the commercial process, including licensing agreements and ongoing purchases. You can describe the feed you need to SourceX in terms of the data, not the businesses that hold it. More delivery patterns are collected in the dataset delivery formats and schemas hub and the broader AI data buyer guides.

Turn your feed specification into a sourcing request

SourceX looks for US businesses that hold the data you describe, reviews rights for each dataset, and delivers it under a license that defines records, uses, term and delivery. Nothing is contracted until a supplier agrees, and a request does not guarantee a match. Start a buyer request with SourceX.

Sources

  1. JSON Schema (json-schema.org), "JSON Schema Specification". https://json-schema.org/specification
  2. Apache Iceberg, "Iceberg schema evolution (docs/evolution.md, 1.1.x)". https://apache.googlesource.com/iceberg/+show/refs/heads/1.1.x/docs/evolution.md
  3. Microsoft Learn, "Detect and manage schema drift". https://learn.microsoft.com/en-us/training/modules/implement-manage-data-quality-constraints-unity-catalog/4-detect-manage-schema-drift
  4. OpenLineage (LF AI & Data), "OpenLineage specification examples". https://openlineage.io/docs/spec/examples
  5. arXiv (MLCommons Croissant authors), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data