Procurement, samples and ongoing supply
Ongoing Data Supply Agreements: Structuring Refresh Deliveries
Quick answer
An ongoing data supply agreement commits a supplier to deliver newly created records on a schedule, under one license and one specification, instead of renegotiating each purchase. For AI buyers it has to settle six terms: the collection window and delivery date of each period, a volume floor and ceiling with a shortfall remedy, the price per period and when it may change, a frozen specification with change control, acceptance of every batch, and how the final delivery and your rights wind down.
By SourceX Editorial · Updated
When recurring supply beats a series of one-off licenses
Choose an ongoing agreement when the value is in records that do not exist yet: a support model that must learn next quarter's products, a retrieval index that must reflect current policies, or an evaluation set that must stay unseen. Repeated one-off purchases repeat rights review and negotiation, and let fields drift between them. Some buyers call it a "forward flow" agreement, a term borrowed from credit markets.
Fresh records matter for three reasons:
- Distribution shift. The WILDS benchmark found that standard training performs substantially worse out of distribution than in distribution across its datasets [1]. New records carry products, codes and phrasing that older data lacks.
- Evaluation contamination. LiveBench's authors note that test-set contamination can quickly make benchmarks obsolete, and they limit it with frequently updated questions from recent math competitions, arXiv papers and news articles [2]. Private evaluation sets age too.
- Market practice. OpenAI's May 2024 arrangement with Reddit was reported to include access to real-time, structured Reddit content [3], and a 2025 Digiday report describes publishers moving from one-time training payments toward usage-based "grounding" deals for retrieval [4].
| Application | What each period delivers | What sets the cadence | What to specify |
|---|---|---|---|
| Continual fine-tuning | New resolved threads, closed tickets, finished workflows | Retraining schedule; pace of product change | Eligibility event; deduplication against earlier batches |
| Retrieval (RAG) refresh | New and changed documents, plus withdrawals | How stale an answer can be before it misleads | Update and delete operations; rights to display retrieved text |
| Evaluation refresh | A held-out slice of new records | Model release cadence; contamination risk | Exclusion from training batches; who else receives the slice |
For cadence from the supplier side, see SourceX's answer on how often AI buyers want fresh data; for text, recurring text data feeds; for earlier stages, the AI training data procurement guide.
Define each period by when records were created, not when they were exported
Define every period by an event in the record's own lifecycle, then set an extraction cut-off and a delivery date, because operational records keep changing after they first appear. A ticket opened on 28 March and resolved on 9 April falls into one quarter only if the agreement names the event, and its satisfaction score may arrive a week later still.
The schedule should state:
- Eligibility event per record type: a support ticket's
resolved_at, a CRM opportunity'sclosed_at, a Jira issue's transition to Done, a call'sended_at, all in UTC. - Settling lag: how long after period close the supplier waits to extract, so late fields (satisfaction scores, refunds, reopen flags) are populated.
- Delivery date: a fixed number of business days after the cut-off; slips fall under the service levels for recurring deliveries.
- Records that change after delivery: resend them as updates under a stable record ID, or freeze each record at first delivery; incremental deliveries versus full refreshes covers signaling inserts, updates and deletes.
The delivery mechanism changes what "delivered" means. Delta Sharing exposes data as share, schema and table, authenticates with a bearer-token profile that can expire, and supports time-travel queries [5]; Snowflake Secure Data Sharing copies no data between accounts and leaves shared objects read-only for the consumer [6]. A share you query in place is not a copy you hold, so define each batch as a named table version and state whether you may copy it. Pulling files from a supplier's Requester Pays bucket puts request and download costs on you every period, while the owner pays storage [7].
Volume bands: commit to a floor and ceiling of usable records
Commit to volume bands rather than exact counts, because operational records arrive at the rate the supplier's business produces them, and measure each band in usable records after de-identification, deduplication and acceptance. A fixed count invites padding in a slow quarter.
Each band needs:
- Floor: the minimum accepted records per period, with a shortfall remedy such as a pro-rata fee reduction, make-up in the next period, or a termination right after consecutive shortfalls.
- Target: the volume the price assumes.
- Ceiling: above it, the buyer may decline surplus or take it at the band price.
- Carry-over: whether surplus in one period offsets a later shortfall.
- Composition caps: a maximum share from any one product line, channel or customer segment, so volume is not met with the easiest records.
Count only after deduplication against every earlier delivery. Near-duplicates are common: one study found a single sentence repeated more than 60,000 times in the C4 corpus, and deduplicated training cut memorized output about tenfold [8]. In a recurring feed, a reopened ticket or a re-exported thread can reappear in a later batch.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Quarter | Rows delivered | Failed schema or de-identification checks | Duplicates of earlier batches | Usable records | Against a 40,000 floor and 50,000 target |
|---|---|---|---|---|---|
| Q3 | 47,000 | 3,500 | 1,800 | 41,700 | Above floor; billed on 41,700 |
| Q4 | 39,000 | 2,100 | 900 | 36,000 | 4,000 short; make-up due next quarter |
Pricing each period, and the points where price may change
Price the unit you accept in each period, and name the dates and events that may reopen price, so neither side can reprice at will mid-term.
| Pricing model | How a period is billed | Fits | Watch for |
|---|---|---|---|
| Per accepted record within band | Accepted records × unit price, floor to ceiling | Volumes that vary by season | Disputes over what counts as accepted |
| Flat period fee with band | Fixed fee while accepted volume stays in band; adjustment outside it | Budget certainty | Paying a full fee for a thin period |
| Tiered unit price | Lower unit price above a threshold | Supply expected to grow | Whether tiers reset each period or accumulate |
| Usage-based | Fees tied to retrieval or display of content [4] | RAG with displayed text | Metering method and audit rights |
Set review points in advance: each anniversary, a breaking specification change, a change in licensed uses, or a band missed for a stated number of periods. Trigger each invoice on written acceptance of that period's batch; milestone payments tied to acceptance covers holdbacks, and cost per usable record makes offers comparable. SourceX does not publish prices; terms depend on scope, volume, history, rights and exclusivity, and are agreed per deal in writing.
Freezing the specification, and the changes that need notice
Attach a versioned data dictionary and preparation specification to the agreement, and classify every change by its effect on your pipeline, with notice and parallel delivery for anything that breaks it. A renamed field breaks ingestion at once; a re-mapped code list or a new redaction tool passes schema checks and quietly changes what the model learns.
| Change type | Example in operational records | What the agreement should require |
|---|---|---|
| Additive | New nullable channel column | Written notice and an updated dictionary |
| Breaking structure | resolution_code renamed close_reason; created_at moves from local-time text to a UTC timestamp | Advance notice; one period delivered in both versions; buyer sign-off |
| Semantic | Support team merges five resolution codes into two | Old-to-new mapping table; flag in the batch manifest; right to re-baseline |
| Preparation | New redaction tool or entity list; placeholders change from <PERSON> to [NAME] | Repeat of the post-processing sample check; tool and version recorded per batch |
| Source system | Supplier migrates from Zendesk to Salesforce Service Cloud | Treat as a new dataset: new sample, rights check and acceptance |
| Collection terms | Supplier revises its customer terms or privacy notice | Notice version in force for each batch's collection window |
Presidio's maintainers caution that, because the tool relies on trained ML models, there is no guarantee it finds all sensitive information [9], so a tool or model swap changes both residual risk and the text the model sees. Collection terms matter too: a February 2024 FTC staff post warned that adopting more permissive data practices, such as AI training or sharing with third parties, through a surreptitious, retroactive change to terms or privacy policies may be unfair or deceptive [10].
Report quality per batch with named measures; ISO/IEC 5259-2 defines a data quality model, measures and guidance on reporting them for analytics and ML data [11]. See handling schema changes across recurring deliveries, data contracts for recurring deliveries and the data dictionary template.
Accepting every batch against the same written tests
Each period's delivery should pass the same written acceptance tests within a fixed inspection window, with rejection, cure and replacement rules that apply per batch; accepting the first delivery says nothing about the twelfth. One annotation vendor advises stating how each term is measured and what remedy applies [12].
Per-batch acceptance checklist:
- Manifest present: period, eligibility event, record count, schema and dictionary versions, preparation method and tool version, notice version, withdrawals applied.
- Usable record count inside the band.
- Schema matches the frozen version, with no unannounced fields.
- Critical-field fill rates within tolerance of the trailing baseline.
- Category mix checked against earlier periods; jumps in one code or channel explained.
- Near-duplicate rate against all earlier deliveries below the agreed threshold.
- Evaluation slice kept out of training batches.
- De-identification sample check run on this batch, not inherited from the first.
- Machine-readable metadata updated, such as Croissant JSON-LD describing dataset metadata, files and record structure [13].
- Written acceptance or rejection inside the window; no deemed acceptance on silence.
For thresholds and cure periods, see acceptance criteria for licensed training data and the dataset acceptance testing process; for replacement records and credits, remedies for defective deliveries; for trends, a supplier performance scorecard and SourceX's guide to data supplier SLAs.
Rights and records that must carry across every batch
Write one grant that covers each future batch on the same terms, and keep a per-batch record of collection period, source and preparation, because rights questions and disclosure duties attach to individual deliveries.
- Grant. License each batch on delivery for the same permitted uses, and state what you keep in each batch after the agreement ends (subscription data licenses for refreshed data, master license agreements with order forms).
- Deidentified status. Under California Civil Code 1798.140(m), information counts as "deidentified" only if the business holding it, among other conditions, contractually obligates recipients to comply with the definition [14]. That obligation travels with every batch.
- Withdrawals. Specify a withdrawal list (record ID, reason code, effective date) and whether withdrawn records must leave training shards, retrieval indexes and evaluation sets; see propagating deletions and corrections.
- Your disclosures. As of October 2026, California AB 2013 requires developers of generative AI systems offered to Californians to post training-data documentation, including dataset sources and collection time periods, by 1 January 2026 and before each later release of a covered system or substantial modification [15]. Per-batch manifests make that update possible when a refresh-trained model ships.
- Diligence. One law firm frames data vendor diligence as a process, not an event [16]; re-attest rights at each anniversary (ongoing due diligence of data vendors).
If you would rather describe the records and cadence than find a supplier yourself, SourceX works with AI data buyers on licensing agreements and ongoing purchases of operational data from US companies. Every dataset goes through rights review, every release is approved by the supplying company, and the license defines which records are included, their permitted uses, how long the license runs and how delivery happens.
Notice, the final period and what you keep at exit
Agree the exit mechanics at signature: the notice period, whether the last period is delivered in full, final acceptance, and what happens to delivered batches, live shares and trained models after the term. Exiting a data contract covers deletion certification; the renewal decision covers notice windows and auto-renewal.
- Notice in periods, not days, so the last batch covers a whole collection window.
- Termination for cause tied to supply failures: consecutive floor shortfalls, an unannounced breaking change, a failed de-identification check, or a rights defect in a batch.
- Final period delivered and accepted under the same tests, with final payment on final acceptance.
- Post-term position: which batches you keep and for what uses, copies taken from shares, and models trained during the license.
A supply schedule to attach to the license
A one-page schedule turns these terms into fields both sides check every period.
Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.
supply_schedule:
dataset: "Resolved support threads, US B2B software supplier"
license_ref: "master license, schedule 3"
term: {periods: 8, period_length: "calendar quarter"}
eligibility:
record_unit: "ticket thread"
event: "resolved_at within the period (UTC)"
settling_lag_days: 14 # wait for late CSAT and reopen flags
changed_after_delivery: "resend as update with the same record_id"
delivery:
due: "15 business days after extraction cut-off"
method: "versioned table in a Delta share; buyer may copy each accepted version"
files: "Parquet plus manifest.json per period"
volume_band_usable_records:
floor: 40000
target: 50000
ceiling: 60000
above_ceiling: "buyer option at band price"
shortfall: "make-up next period; termination right after 2 consecutive shortfalls"
composition_cap: "no product line above 40% of a period"
pricing:
unit: "accepted record"
review_points: ["anniversary", "breaking spec change", "change in licensed uses"]
invoice_trigger: "written acceptance of the period's batch"
specification:
data_dictionary_version: "2.0"
preparation_spec_version: "1.3"
breaking_change_notice: "one full period"
parallel_delivery: "one period in old and new schema"
acceptance:
inspection_window_business_days: 10
tests: ["band count", "schema", "fill rates", "cross-batch duplicates", "de-identification sample", "manifest"]
deemed_acceptance_on_silence: false
withdrawals: {format: "record_id, reason_code, effective_date", purge: ["training shards", "retrieval index", "eval sets"]}
exit:
notice: "one full period"
final_period: "delivered and accepted in full"
post_term: "per license sections on retained batches and trained models"
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Planning recurring deliveries of new records?
Describe the records you need, the cadence, the volume band and the uses each batch must support. SourceX looks for US businesses that hold that data, checks the supplier's licensing permissions, and manages the license, delivery and future purchases; nothing is contracted until a supplier agrees. Describe your recurring data needs.
Sources
- Koh et al., "WILDS: A Benchmark of in-the-Wild Distribution Shifts" (2020). https://arxiv.org/pdf/2012.07421
- White et al., "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (ICLR 2025). https://www.arxiv.org/pdf/2406.19314
- TechCrunch, "OpenAI inks deal to train AI on Reddit data" (2024; descriptive title from URL). https://techcrunch.com/2024/05/16/openai-inks-deal-to-train-ai-on-reddit-data
- Digiday, "WTF is AI 'grounding' licensing, and why do publishers say it matters over training deals?" (2025). https://digiday.com/media/wtf-is-ai-grounding-licensing-and-why-do-publishers-say-it-matters-over-training-deals/
- delta-io/delta-sharing documentation (Mintlify rendering), "Delta Sharing protocol: REST APIs". https://www.mintlify.com/delta-io/delta-sharing/protocol/rest-apis
- Snowflake Documentation, "About Secure Data Sharing". https://docs.snowflake.com/en/user-guide/data-sharing-intro.html
- Amazon Web Services, "Using Requester Pays general purpose buckets for storage transfers and usage". https://docs.aws.amazon.com/AmazonS3/latest/dev/RequesterPaysBuckets.html
- Lee et al., "Deduplicating Training Data Makes Language Models Better" (ACL 2022). https://arxiv.org/abs/2107.06499v1
- Microsoft presidio project, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- Federal Trade Commission, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- Digital Divide Data, "Data annotation service agreements: measurement and remedies" (vendor blog; descriptive title). https://www.digitaldividedata.com/?p=24243
- Akhtar et al., "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- Foley & Lardner LLP, "Data vendor due diligence commentary" (descriptive title). https://foley.com/?p=49353
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.