Procurement, samples and ongoing supply
SLAs for Recurring Data Deliveries: Freshness, Latency and Completeness
Quick answer
A data feed SLA for recurring training data should measure five things per batch: freshness lag (time from record creation to delivery), delivery date tolerance (how late a batch may land), completeness (records and required fields against a control total), schema stability (no unannounced breaking changes) and validity (share of records passing agreed rules). Each metric needs a written formula, a measurement point, a threshold, a monthly report and a remedy for repeated misses, otherwise the SLA cannot be enforced.
By SourceX Editorial · Updated
This guide is for procurement managers running an ongoing supply contract for continual fine-tuning, RAG corpora or evaluation refreshes. It assumes the commercial structure is already settled (see ongoing data supply agreements) and focuses on the service levels themselves. For one-off delivery terms, the owner page on data supplier SLAs covers the broader ground.
Why recurring feeds need different metrics than a one-off delivery
A one-off delivery is judged once at acceptance; a recurring feed is judged every period, so its SLA must define how each batch is measured and how misses accumulate. Acceptance criteria for a single file (see acceptance criteria for licensed training data) still apply per batch, but they say nothing about timing, drift or trend. The commercial risk in a feed is gradual degradation: batches arrive a little later, a field quietly empties, a category stops appearing.
Timeliness matters most when the data feeds a retrieval index or a refreshed eval set. Research on question answering under temporal conflict shows that models struggle when knowledge evolves over time [2], which is why RAG and continual-update teams care about the age of what they receive, not only its volume. If your team is still deciding how often it needs data, the SourceX answer on how often AI buyers want fresh data is a useful starting point.
The core metrics and how to define each one
Every metric in a feed SLA should be stated as a formula over fields that exist in the delivered data or the batch manifest. ISO/IEC 5259-2 provides a vocabulary of measurable data quality characteristics for ML data, including completeness, accuracy, consistency and timeliness measures, and it is a sensible reference when you need neutral definitions both parties accept [1]. The definitions below translate those ideas into contract terms.
- Freshness lag. For each record,
delivered_at - source_created_at(orsource_updated_atfor changed records). Commit to a percentile, such as the 95th, not an average, because a few very old records can hide behind a good mean. Require the supplier to populate the source timestamp; without it, freshness cannot be measured. - Delivery date tolerance. The batch is on time if its manifest lands in the agreed location within a window after the scheduled cut-off. Measure from the timestamp of the completed upload or share, not the notification email.
- Volume and record completeness. Delivered record count divided by the supplier's control total for the period (the count its source system reports for the same filter). Pair it with an expected range so a batch that is "complete" against a broken control total still gets flagged.
- Field completeness. Fill rate for each required field (
ticket_id,created_at,resolution_code, and so on), measured against the agreed data dictionary. - Schema stability. Zero breaking changes (renamed, removed or retyped columns, changed enumerations) without the agreed notice period; additive columns allowed with notice.
- Validity. Share of records passing agreed rules: parseable timestamps, enumerations within the dictionary, no duplicate primary keys, redaction tokens in the expected format.
- Correction turnaround. Time from your defect notice to a corrected or replacement batch.
Data services vendors describe turnaround and capacity commitments with stated measurement terms [3]; a common gap in feed agreements is a commitment with no definition of where and how it is measured.
Freshness lag versus delivery latency
Freshness lag measures the age of the records; delivery latency measures whether the batch arrived on schedule, and a supplier can meet one while failing the other. A monthly batch delivered on the due date can still contain records that are 60 days old if the supplier's extraction cut-off sits a month earlier. Conversely, a batch of very recent records can land three days late.
Write both into the SLA. The refresh pattern also changes what freshness means: in a full refresh the whole dataset is replaced, while an incremental delivery moves only new or changed records [4]. For incremental feeds, measure freshness on the delta and add a separate metric for late-arriving updates to records you already hold. The tradeoffs between the two patterns are covered in incremental deliveries vs full refreshes.
Completeness against a control total
Completeness is only measurable when the supplier delivers a control total produced independently of the export job. A record count taken from the exported file proves nothing about what the source system held. Ask for counts from the source system's own reporting for the same date range and filters, broken down by the strata you care about (product line, region, ticket category, document type).
Completeness should also be measured after the supplier's privacy processing. When names, emails and account numbers are removed or replaced, some suppliers drop records rather than redact them, which shows up as a quiet volume shortfall in specific strata. Require the manifest to report records excluded at each preparation step and the reason code, so you can tell a privacy exclusion from an extraction failure.
Schema stability and change notice
Schema stability is a service level on its own because one unannounced column rename can break every downstream job that reads the feed. Ingestion pipelines need an explicit policy for drift (fail, quarantine, or evolve the schema) [5], and the contract should match that policy: breaking changes require advance written notice, a versioned data dictionary, and a parallel period in which old and new layouts are both delivered.
Machine-readable metadata makes this checkable. A Croissant description, a JSON-LD vocabulary built on schema.org that records dataset metadata, file resources and record structure [6], or a plain versioned JSON schema shipped with each batch, lets your pipeline compare the declared structure to the previous release automatically. The mechanics of handling changes once they happen belong to schema evolution across recurring deliveries; the field-level specification itself is covered in data contracts for recurring deliveries.
Where and how each metric is measured
Each metric needs a named measurement point, because disputes about late or incomplete batches usually turn on whose clock counts. For delivery tolerance, use the completion timestamp in the shared storage location or the table version timestamp in a share. Open protocols such as Delta Sharing expose tables over REST on top of Parquet files in cloud object storage [7], which gives both parties the same version history to read from.
For record-level metrics, agree on who runs the checks. A common split is that the supplier runs validation before release and includes results in the manifest, and the buyer re-runs the same rule set on arrival; where the two disagree, the buyer's run on the delivered files prevails unless the supplier shows a transfer error. Store rule definitions in version control and reference the version in each report.
SLA schedule and monthly report template
The SLA schedule below shows how to express each metric with a formula, threshold and remedy so a reviewer can check compliance from the delivered data alone.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Metric | Formula | Measured at | Threshold (example) | Remedy trigger |
|---|---|---|---|---|
| Freshness lag | p95 of delivered_at - source_created_at | Buyer on arrival | 14 days or less | Two consecutive misses |
| Delivery tolerance | Upload complete time minus scheduled cut-off | Storage or share version timestamp | 2 business days | Any batch more than 5 business days late |
| Record completeness | Delivered records / source control total | Manifest vs buyer count | 98% or more per stratum | Any stratum below 95% |
| Field completeness | Non-null required fields / records | Buyer validation run | 99% per required field | Two misses in a quarter |
| Schema stability | Breaking changes without notice | Dictionary diff | Zero | Any occurrence |
| Validity | Records passing rule set v-n / records | Both parties, buyer prevails | 99.5% or more | Two misses in a quarter |
| Correction turnaround | Replacement batch time minus defect notice | Ticket log | 10 business days | Any miss |
A monthly SLA report from the supplier can be as simple as one record per batch:
Illustrative example: invented to show structure; it does not describe an available dataset.
batch_id: support-tickets-2026-09
scheduled_cutoff: 2026-09-30
delivered_at: 2026-10-02T14:05:00Z
schema_version: 3.2.0
schema_change: additive # new column: escalation_tier, notice sent 2026-08-29
records_delivered: 48210
source_control_total: 48903
excluded_by_step:
deidentification: 512
dedup: 181
freshness_lag_p95_days: 11
required_field_fill:
ticket_id: 1.000
created_at: 1.000
resolution_code: 0.987
validation_rule_set: v7
validity_rate: 0.996
open_defects: 0
Late data feed penalties and service credits
Remedies for a feed should escalate with repetition rather than punish a single late batch, because the buyer's real loss is a pattern that stalls retraining or eval refreshes. A practical ladder is: a corrective action plan after the first trigger, a service credit against the next period's fee after repeated misses in a rolling window, and a right to suspend or terminate the affected stream after sustained failure. Tie credits to the batch's share of the period fee so they stay proportionate.
Two cautions apply. First, exclude misses caused by the buyer, such as a late change request or an unavailable receiving location, and define those exclusions precisely. Second, credits are not a quality fix; keep the right to reject a batch that fails acceptance and receive a replacement, as described in acceptance sampling for dataset deliveries.
How SourceX approaches recurring purchases
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Each dataset is rights-reviewed and delivered under a license that defines the records, uses, term and delivery, and every release is approved by the supplying company. Delivery runs through private, access-controlled workflows only after an executed agreement. If you are scoping a refresh feed, describe the data you need to SourceX, and see the procurement hub for related guides.
Set service levels for your next data refresh feed
Describe the recurring data you need, including cadence, fields and freshness expectations, and SourceX will look for US businesses that hold it; a request does not guarantee a match, and nothing is contracted until a supplier agrees. Pricing and allowed uses are agreed per deal in a license. Start a buyer request at sourcex.si/buyers.
Frequently asked questions
Should freshness be an average or a percentile?
Use a percentile, typically the 95th, plus a hard maximum for any single record. Averages let a small block of stale records through, and stale records concentrated in one stratum can skew a fine-tuning mix or leave a retrieval index with outdated answers.
What if the supplier's source system has no reliable creation timestamp?
Agree on a proxy field in the data dictionary (last-modified time, posting date, closed date) and state in the SLA which one is used. Without a defined timestamp, freshness lag is not measurable and should be dropped rather than estimated.
How often should SLA performance be reviewed?
Report per batch and review monthly or quarterly depending on cadence. A quarterly review that looks at trends across batches, such as a slowly falling fill rate on one field, catches degradation that per-batch thresholds miss.
Sources
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- arXiv, "Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs" (2025). https://arxiv.org/html/2506.07270v1
- Digital Divide Data, "Dataset acceptance criteria and SLAs for annotated training data" (2024). https://www.digitaldividedata.com/?p=24243
- Airbyte, "Full Refresh vs Incremental Refresh in ETL: How to Decide?". https://airbyte.com/data-engineering-resources/full-refresh-vs-incremental-refresh
- Microsoft Learn, "Detect and manage schema drift". https://learn.microsoft.com/en-us/training/modules/implement-manage-data-quality-constraints-unity-catalog/4-detect-manage-schema-drift
- MLCommons (Akhtar et al.), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- Databricks, "Introducing Delta Sharing: An Open Protocol for Secure Data Sharing" (2021). https://www.databricks.com/blog/2021/05/26/introducing-delta-sharing-an-open-protocol-for-secure-data-sharing.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.