Procurement, samples and ongoing supply
Paying for Data in Milestones Tied to Acceptance
Quick answer
Milestone payments for data delivery split the license fee across verifiable events (signature, sample acceptance, acceptance of each delivery tranche and final acceptance) and keep a holdback for defects that only show up once the data is in a training pipeline. Each payment is released by a written acceptance notice issued against criteria agreed before delivery, not by a delivery date. This caps the amount you have at risk at any point. It also gives the supplier a predictable path to getting paid.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why prepayment is the expensive failure mode for licensed datasets
Paying the full fee up front moves all defect risk to the buyer, because a dataset's flaws are usually invisible until someone parses, profiles and trains on it. Practitioners who audit robot training datasets frame acceptance as a verdict reached before payment, on the view that a problem found in a sample is cheaper to resolve than one found after the invoice has cleared [1]. After full payment your remaining leverage is a contractual remedy claim, which is slower and costlier than withholding the next payment.
Data has three traits that make prepayment especially risky compared with buying software seats:
- Defects are statistical. Label error rates, duplicate rates and null rates are properties of the whole delivery, so a clean sample does not prove a clean corpus.
- Some defects are latent. Residual personal data, near-duplicates of your eval set or a skewed class distribution may only appear during de-duplication, contamination checks or the first fine-tuning run.
- Deliveries arrive in pieces. Large or ongoing purchases ship as tranches (by month of records, by business unit or by modality), and each tranche can differ in quality.
Milestones align cash with those realities. The structure should be settled during commercial negotiation, alongside the acceptance criteria for licensed training data, because a payment schedule without measurable criteria simply moves the dispute to a later date.
The four milestones that map to how data is actually verified
A workable schedule ties each payment to a gate you can verify yourself: execution, sample acceptance, tranche acceptance and final acceptance with a holdback. Each gate needs an objective trigger, an owner on your side and a deemed-acceptance rule so the supplier is not left waiting indefinitely.
- Signature (execution) payment. A modest portion compensates the supplier for preparation work such as extraction, de-identification and schema mapping. Keep it proportionate to that work, because it is the one payment you make before seeing production data.
- Sample acceptance. The supplier delivers a representative sample in the production schema and format. Use the process in how to request a training data sample and score it against the same thresholds you will apply to the full delivery.
- Tranche acceptance. Each delivery batch passes structural checks (manifest completeness, checksums, schema conformance) and a statistical sample inspection before its payment is released.
- Final acceptance and holdback release. After the last tranche, a retained portion is held through a defined post-acceptance window for latent defects, then released automatically unless you have issued a valid defect notice.
The supplier-facing view of the same mechanics, written for companies being paid, lives in the guide on buyer acceptance and getting paid.
Illustrative milestone schedule for a tranche-delivered dataset
The schedule below shows the structure, not a recommended split; your percentages depend on preparation cost, tranche count and how much of the value sits in the later data.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Milestone | Trigger (objective evidence) | Illustrative share | Acceptance window | Deemed acceptance |
|---|---|---|---|---|
| M0 Execution | License signed by both parties | 10% | n/a | n/a |
| M1 Sample | Sample of 2,000 records in production schema passes thresholds | 15% | 10 business days | Yes, if no rejection notice |
| M2 Tranche 1 | Records Jan-Jun: manifest, SHA-256 checksums and schema checks pass; sampling plan accepts | 25% | 15 business days | Yes |
| M3 Tranche 2 | Records Jul-Dec: same checks | 25% | 15 business days | Yes |
| M4 Final | All tranches accepted; documentation and data card delivered | 15% | 10 business days | Yes |
| H Holdback | No unresolved latent-defect notice at end of window | 10% | 60 days after M4 | Released automatically |
Read the table as a set of linked decisions. The holdback window should be long enough to cover your real pipeline (de-duplication, contamination checks against evals, a first training run), but bounded, so the supplier can book revenue.
Writing triggers that finance and engineering both accept
A payment trigger works only if accounts payable can verify it from a document and engineering can produce that document from a check it already runs. The cleanest trigger is a dated acceptance certificate that references the tranche manifest hash and the results of named tests.
Structural checks are binary and should come first. A tranche manifest lists every file with its byte size and checksum, and your intake job recomputes them; see dataset manifests and checksums. Format checks catch truncation early: a Parquet file must begin and end with the PAR1 magic number and carries its metadata in the footer, so a cut-off upload fails before anyone profiles a column [6].
Content checks are statistical and should use a declared sampling plan rather than an engineer's impression. A single sampling plan is defined by a sample size n and an acceptance number c, chosen against a producer's risk point (AQL) and a consumer's risk point (LTPD) [2]. Tabled schemes in the MIL-STD-105 lineage index n and c by lot size, inspection level and AQL [3], and ANSI/ASQ Z1.4 is built for a continuing stream of lots with switching between normal, tightened and reduced inspection [4], which maps naturally onto tranches. The acceptance sampling guide for dataset deliveries adapts these plans to record and label defects.
Good triggers name the following elements:
- the tranche identifier and manifest hash being accepted;
- each test, its threshold and the sampling plan (n, c, defect definitions);
- the acceptance window in business days and the deemed-acceptance rule;
- who signs the certificate on the buyer side and where it is sent.
Sizing the holdback for latent defects
A holdback is the retained share of the fee that covers defects your acceptance tests could not reasonably detect within the acceptance window. Size it to the cost of the most likely latent defect, not as a penalty, and define precisely which defects it covers.
Residual personal data is the textbook latent defect. Automated de-identification tools rely on trained detectors, and the Presidio project itself cautions that there is no guarantee it will find all sensitive information [5]. A finding from a later, larger scan should therefore trigger a cure obligation and, if uncured, a claim on the holdback; the residual PII audit sampling guide covers how to size those scans.
Other defects that commonly justify a holdback:
- Duplicate or near-duplicate records found by MinHash or embedding de-duplication across tranches.
- Eval contamination, where delivered records overlap your held-out test sets.
- Documentation gaps, such as a missing field dictionary or provenance notes your counsel needs for training data disclosures.
- Distribution drift between tranches, for example a category whose share collapses in the second half of the delivery.
What should not draw on the holdback: disappointment with model gains that were never specified as acceptance criteria. If downstream lift matters, define it as a measurable test in the license, as discussed in estimating a dataset's value before you buy.
Rejection, cure and what happens to the payment
A rejected milestone should pause that payment, start a cure period and lead to one of three outcomes: re-delivery that passes, a partial acceptance at an adjusted amount, or termination of the undelivered scope. Write those outcomes down before signing so no one improvises them under deadline pressure.
The mechanics of notices, inspection windows and cure periods are covered in dataset acceptance testing. The payment-side questions to settle are narrower:
- Does a rejected tranche delay later tranches, or can they be accepted independently?
- Is partial acceptance priced per accepted record, per accepted file or as a negotiated credit?
- After a failed cure, is the remedy replacement records, a credit against the next payment or a refund? See remedies when a data delivery fails.
- Does the license or usage right for a rejected tranche lapse, and must you delete it?
Answering these up front keeps the milestone schedule enforceable.
Adapting milestones to ongoing supply and refresh deliveries
For ongoing purchases, the milestone pattern becomes a recurring cycle: each refresh is a tranche with its own acceptance window, payment and, optionally, a rolling holdback. Recurring deliveries benefit most from switching rules, because a supplier with a clean track record can move to reduced inspection, while a failed refresh moves the next one to tightened inspection [4].
Keep three things stable across cycles: the schema contract, the test suite and the certificate template. Changes to any of them should be versioned amendments, not email agreements. The ongoing data supply agreements guide covers refresh cadence and scope, and incremental versus full refresh deliveries affects what each acceptance cycle must re-test.
When comparing suppliers, model the payment schedule into effective cost. Two quotes with the same headline price differ in risk if one requires most of the fee at signature; comparing data vendor quotes shows how to normalize to cost per usable record.
Pre-signature checklist for a milestone payment schedule
Use this checklist before the license goes to signature; every item should be answerable with a clause reference.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Each milestone has one objective trigger and a named evidence document.
- Acceptance criteria and sampling plans are attached as a schedule, not referenced informally.
- Acceptance windows and deemed-acceptance rules are stated for every milestone.
- The holdback amount, window, covered defects and automatic release are defined.
- Rejection, cure period, partial acceptance and termination outcomes are tied to payments.
- Invoices may only issue after a signed acceptance certificate for that milestone.
- Refresh cycles reuse the same test suite and certificate template, with versioned changes.
For how payments typically move once terms are agreed, see how data licensing payments are made, and for the wider purchasing sequence, the AI training data procurement hub.
Where SourceX fits in a licensed data purchase
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Pricing and allowed uses are agreed in a license that defines records, uses, term and delivery, and nothing is contracted until a supplier agrees. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. If you are planning a staged purchase, you can describe the data you need to SourceX.
Plan a staged data license with SourceX
SourceX does not hold datasets in stock: it looks for US businesses that hold the data you describe, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, with diligence materials prepared per dataset, and terms are agreed per deal. Tell SourceX what data you are looking for.
Sources
- Contra (independent practitioner listing), "Robot dataset acceptance audit: a verdict before you pay". https://contra.com/s/ee6vvEQU-robot-dataset-acceptance-audit-a-verdict-before-you-pay
- National Institute of Standards and Technology, "NIST/SEMATECH e-Handbook of Statistical Methods: How do you choose a single sampling plan?". https://itl.nist.gov/div898/handbook/pmc/section2/pmc23.htm
- National Institute of Standards and Technology, "NIST/SEMATECH e-Handbook of Statistical Methods: Choosing a sampling plan: MIL Std 105D". https://itl.nist.gov/div898/handbook/pmc/section2/pmc231.htm
- ASQ Quality Press, "ASQ/ANSI Z1.4:2003 (R2018): Sampling Procedures and Tables for Inspection by Attributes". https://asq.org/quality-press/display-item?item=T1164
- Microsoft (microsoft/presidio project), indexed on pkg.go.dev, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- The Apache Software Foundation (Apache Parquet project), "File Format". https://parquet.apache.org/docs/file-format/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.