Skip to content

Procurement, samples and ongoing supply

Ongoing Data Supply Agreements: Structuring Refresh Deliveries

Quick answer

An ongoing data supply agreement commits a supplier to deliver newly created records on a schedule, under one license and one specification, instead of renegotiating each purchase. For AI buyers it has to settle six terms: the collection window and delivery date of each period, a volume floor and ceiling with a shortfall remedy, the price per period and when it may change, a frozen specification with change control, acceptance of every batch, and how the final delivery and your rights wind down.

By SourceX Editorial · Updated

When recurring supply beats a series of one-off licenses

Choose an ongoing agreement when the value is in records that do not exist yet: a support model that must learn next quarter's products, a retrieval index that must reflect current policies, or an evaluation set that must stay unseen. Repeated one-off purchases repeat rights review and negotiation, and let fields drift between them. Some buyers call it a "forward flow" agreement, a term borrowed from credit markets.

Fresh records matter for three reasons:

  • Distribution shift. The WILDS benchmark found that standard training performs substantially worse out of distribution than in distribution across its datasets [1]. New records carry products, codes and phrasing that older data lacks.
  • Evaluation contamination. LiveBench's authors note that test-set contamination can quickly make benchmarks obsolete, and they limit it with frequently updated questions from recent math competitions, arXiv papers and news articles [2]. Private evaluation sets age too.
  • Market practice. OpenAI's May 2024 arrangement with Reddit was reported to include access to real-time, structured Reddit content [3], and a 2025 Digiday report describes publishers moving from one-time training payments toward usage-based "grounding" deals for retrieval [4].
ApplicationWhat each period deliversWhat sets the cadenceWhat to specify
Continual fine-tuningNew resolved threads, closed tickets, finished workflowsRetraining schedule; pace of product changeEligibility event; deduplication against earlier batches
Retrieval (RAG) refreshNew and changed documents, plus withdrawalsHow stale an answer can be before it misleadsUpdate and delete operations; rights to display retrieved text
Evaluation refreshA held-out slice of new recordsModel release cadence; contamination riskExclusion from training batches; who else receives the slice

For cadence from the supplier side, see SourceX's answer on how often AI buyers want fresh data; for text, recurring text data feeds; for earlier stages, the AI training data procurement guide.

Define each period by when records were created, not when they were exported

Define every period by an event in the record's own lifecycle, then set an extraction cut-off and a delivery date, because operational records keep changing after they first appear. A ticket opened on 28 March and resolved on 9 April falls into one quarter only if the agreement names the event, and its satisfaction score may arrive a week later still.

The schedule should state:

  • Eligibility event per record type: a support ticket's resolved_at, a CRM opportunity's closed_at, a Jira issue's transition to Done, a call's ended_at, all in UTC.
  • Settling lag: how long after period close the supplier waits to extract, so late fields (satisfaction scores, refunds, reopen flags) are populated.
  • Delivery date: a fixed number of business days after the cut-off; slips fall under the service levels for recurring deliveries.
  • Records that change after delivery: resend them as updates under a stable record ID, or freeze each record at first delivery; incremental deliveries versus full refreshes covers signaling inserts, updates and deletes.

The delivery mechanism changes what "delivered" means. Delta Sharing exposes data as share, schema and table, authenticates with a bearer-token profile that can expire, and supports time-travel queries [5]; Snowflake Secure Data Sharing copies no data between accounts and leaves shared objects read-only for the consumer [6]. A share you query in place is not a copy you hold, so define each batch as a named table version and state whether you may copy it. Pulling files from a supplier's Requester Pays bucket puts request and download costs on you every period, while the owner pays storage [7].

Volume bands: commit to a floor and ceiling of usable records

Commit to volume bands rather than exact counts, because operational records arrive at the rate the supplier's business produces them, and measure each band in usable records after de-identification, deduplication and acceptance. A fixed count invites padding in a slow quarter.

Each band needs:

  • Floor: the minimum accepted records per period, with a shortfall remedy such as a pro-rata fee reduction, make-up in the next period, or a termination right after consecutive shortfalls.
  • Target: the volume the price assumes.
  • Ceiling: above it, the buyer may decline surplus or take it at the band price.
  • Carry-over: whether surplus in one period offsets a later shortfall.
  • Composition caps: a maximum share from any one product line, channel or customer segment, so volume is not met with the easiest records.

Count only after deduplication against every earlier delivery. Near-duplicates are common: one study found a single sentence repeated more than 60,000 times in the C4 corpus, and deduplicated training cut memorized output about tenfold [8]. In a recurring feed, a reopened ticket or a re-exported thread can reappear in a later batch.

Illustrative example: invented to show structure; it does not describe an available dataset.

QuarterRows deliveredFailed schema or de-identification checksDuplicates of earlier batchesUsable recordsAgainst a 40,000 floor and 50,000 target
Q347,0003,5001,80041,700Above floor; billed on 41,700
Q439,0002,10090036,0004,000 short; make-up due next quarter

Pricing each period, and the points where price may change

Price the unit you accept in each period, and name the dates and events that may reopen price, so neither side can reprice at will mid-term.

Pricing modelHow a period is billedFitsWatch for
Per accepted record within bandAccepted records × unit price, floor to ceilingVolumes that vary by seasonDisputes over what counts as accepted
Flat period fee with bandFixed fee while accepted volume stays in band; adjustment outside itBudget certaintyPaying a full fee for a thin period
Tiered unit priceLower unit price above a thresholdSupply expected to growWhether tiers reset each period or accumulate
Usage-basedFees tied to retrieval or display of content [4]RAG with displayed textMetering method and audit rights

Set review points in advance: each anniversary, a breaking specification change, a change in licensed uses, or a band missed for a stated number of periods. Trigger each invoice on written acceptance of that period's batch; milestone payments tied to acceptance covers holdbacks, and cost per usable record makes offers comparable. SourceX does not publish prices; terms depend on scope, volume, history, rights and exclusivity, and are agreed per deal in writing.

Freezing the specification, and the changes that need notice

Attach a versioned data dictionary and preparation specification to the agreement, and classify every change by its effect on your pipeline, with notice and parallel delivery for anything that breaks it. A renamed field breaks ingestion at once; a re-mapped code list or a new redaction tool passes schema checks and quietly changes what the model learns.

Change typeExample in operational recordsWhat the agreement should require
AdditiveNew nullable channel columnWritten notice and an updated dictionary
Breaking structureresolution_code renamed close_reason; created_at moves from local-time text to a UTC timestampAdvance notice; one period delivered in both versions; buyer sign-off
SemanticSupport team merges five resolution codes into twoOld-to-new mapping table; flag in the batch manifest; right to re-baseline
PreparationNew redaction tool or entity list; placeholders change from <PERSON> to [NAME]Repeat of the post-processing sample check; tool and version recorded per batch
Source systemSupplier migrates from Zendesk to Salesforce Service CloudTreat as a new dataset: new sample, rights check and acceptance
Collection termsSupplier revises its customer terms or privacy noticeNotice version in force for each batch's collection window

Presidio's maintainers caution that, because the tool relies on trained ML models, there is no guarantee it finds all sensitive information [9], so a tool or model swap changes both residual risk and the text the model sees. Collection terms matter too: a February 2024 FTC staff post warned that adopting more permissive data practices, such as AI training or sharing with third parties, through a surreptitious, retroactive change to terms or privacy policies may be unfair or deceptive [10].

Report quality per batch with named measures; ISO/IEC 5259-2 defines a data quality model, measures and guidance on reporting them for analytics and ML data [11]. See handling schema changes across recurring deliveries, data contracts for recurring deliveries and the data dictionary template.

Accepting every batch against the same written tests

Each period's delivery should pass the same written acceptance tests within a fixed inspection window, with rejection, cure and replacement rules that apply per batch; accepting the first delivery says nothing about the twelfth. One annotation vendor advises stating how each term is measured and what remedy applies [12].

Per-batch acceptance checklist:

  • Manifest present: period, eligibility event, record count, schema and dictionary versions, preparation method and tool version, notice version, withdrawals applied.
  • Usable record count inside the band.
  • Schema matches the frozen version, with no unannounced fields.
  • Critical-field fill rates within tolerance of the trailing baseline.
  • Category mix checked against earlier periods; jumps in one code or channel explained.
  • Near-duplicate rate against all earlier deliveries below the agreed threshold.
  • Evaluation slice kept out of training batches.
  • De-identification sample check run on this batch, not inherited from the first.
  • Machine-readable metadata updated, such as Croissant JSON-LD describing dataset metadata, files and record structure [13].
  • Written acceptance or rejection inside the window; no deemed acceptance on silence.

For thresholds and cure periods, see acceptance criteria for licensed training data and the dataset acceptance testing process; for replacement records and credits, remedies for defective deliveries; for trends, a supplier performance scorecard and SourceX's guide to data supplier SLAs.

Rights and records that must carry across every batch

Write one grant that covers each future batch on the same terms, and keep a per-batch record of collection period, source and preparation, because rights questions and disclosure duties attach to individual deliveries.

  • Grant. License each batch on delivery for the same permitted uses, and state what you keep in each batch after the agreement ends (subscription data licenses for refreshed data, master license agreements with order forms).
  • Deidentified status. Under California Civil Code 1798.140(m), information counts as "deidentified" only if the business holding it, among other conditions, contractually obligates recipients to comply with the definition [14]. That obligation travels with every batch.
  • Withdrawals. Specify a withdrawal list (record ID, reason code, effective date) and whether withdrawn records must leave training shards, retrieval indexes and evaluation sets; see propagating deletions and corrections.
  • Your disclosures. As of October 2026, California AB 2013 requires developers of generative AI systems offered to Californians to post training-data documentation, including dataset sources and collection time periods, by 1 January 2026 and before each later release of a covered system or substantial modification [15]. Per-batch manifests make that update possible when a refresh-trained model ships.
  • Diligence. One law firm frames data vendor diligence as a process, not an event [16]; re-attest rights at each anniversary (ongoing due diligence of data vendors).

If you would rather describe the records and cadence than find a supplier yourself, SourceX works with AI data buyers on licensing agreements and ongoing purchases of operational data from US companies. Every dataset goes through rights review, every release is approved by the supplying company, and the license defines which records are included, their permitted uses, how long the license runs and how delivery happens.

Notice, the final period and what you keep at exit

Agree the exit mechanics at signature: the notice period, whether the last period is delivered in full, final acceptance, and what happens to delivered batches, live shares and trained models after the term. Exiting a data contract covers deletion certification; the renewal decision covers notice windows and auto-renewal.

  • Notice in periods, not days, so the last batch covers a whole collection window.
  • Termination for cause tied to supply failures: consecutive floor shortfalls, an unannounced breaking change, a failed de-identification check, or a rights defect in a batch.
  • Final period delivered and accepted under the same tests, with final payment on final acceptance.
  • Post-term position: which batches you keep and for what uses, copies taken from shares, and models trained during the license.

A supply schedule to attach to the license

A one-page schedule turns these terms into fields both sides check every period.

Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.

supply_schedule:
  dataset: "Resolved support threads, US B2B software supplier"
  license_ref: "master license, schedule 3"
  term: {periods: 8, period_length: "calendar quarter"}
  eligibility:
    record_unit: "ticket thread"
    event: "resolved_at within the period (UTC)"
    settling_lag_days: 14             # wait for late CSAT and reopen flags
    changed_after_delivery: "resend as update with the same record_id"
  delivery:
    due: "15 business days after extraction cut-off"
    method: "versioned table in a Delta share; buyer may copy each accepted version"
    files: "Parquet plus manifest.json per period"
  volume_band_usable_records:
    floor: 40000
    target: 50000
    ceiling: 60000
    above_ceiling: "buyer option at band price"
    shortfall: "make-up next period; termination right after 2 consecutive shortfalls"
    composition_cap: "no product line above 40% of a period"
  pricing:
    unit: "accepted record"
    review_points: ["anniversary", "breaking spec change", "change in licensed uses"]
    invoice_trigger: "written acceptance of the period's batch"
  specification:
    data_dictionary_version: "2.0"
    preparation_spec_version: "1.3"
    breaking_change_notice: "one full period"
    parallel_delivery: "one period in old and new schema"
  acceptance:
    inspection_window_business_days: 10
    tests: ["band count", "schema", "fill rates", "cross-batch duplicates", "de-identification sample", "manifest"]
    deemed_acceptance_on_silence: false
  withdrawals: {format: "record_id, reason_code, effective_date", purge: ["training shards", "retrieval index", "eval sets"]}
  exit:
    notice: "one full period"
    final_period: "delivered and accepted in full"
    post_term: "per license sections on retained batches and trained models"

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Planning recurring deliveries of new records?

Describe the records you need, the cadence, the volume band and the uses each batch must support. SourceX looks for US businesses that hold that data, checks the supplier's licensing permissions, and manages the license, delivery and future purchases; nothing is contracted until a supplier agrees. Describe your recurring data needs.

Sources

  1. Koh et al., "WILDS: A Benchmark of in-the-Wild Distribution Shifts" (2020). https://arxiv.org/pdf/2012.07421
  2. White et al., "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (ICLR 2025). https://www.arxiv.org/pdf/2406.19314
  3. TechCrunch, "OpenAI inks deal to train AI on Reddit data" (2024; descriptive title from URL). https://techcrunch.com/2024/05/16/openai-inks-deal-to-train-ai-on-reddit-data
  4. Digiday, "WTF is AI 'grounding' licensing, and why do publishers say it matters over training deals?" (2025). https://digiday.com/media/wtf-is-ai-grounding-licensing-and-why-do-publishers-say-it-matters-over-training-deals/
  5. delta-io/delta-sharing documentation (Mintlify rendering), "Delta Sharing protocol: REST APIs". https://www.mintlify.com/delta-io/delta-sharing/protocol/rest-apis
  6. Snowflake Documentation, "About Secure Data Sharing". https://docs.snowflake.com/en/user-guide/data-sharing-intro.html
  7. Amazon Web Services, "Using Requester Pays general purpose buckets for storage transfers and usage". https://docs.aws.amazon.com/AmazonS3/latest/dev/RequesterPaysBuckets.html
  8. Lee et al., "Deduplicating Training Data Makes Language Models Better" (ACL 2022). https://arxiv.org/abs/2107.06499v1
  9. Microsoft presidio project, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  10. Federal Trade Commission, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  11. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
  12. Digital Divide Data, "Data annotation service agreements: measurement and remedies" (vendor blog; descriptive title). https://www.digitaldividedata.com/?p=24243
  13. Akhtar et al., "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  14. California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  15. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  16. Foley & Lardner LLP, "Data vendor due diligence commentary" (descriptive title). https://foley.com/?p=49353

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data