Procurement, samples and ongoing supply
Acceptance criteria for training data delivery: thresholds to write into the order
Quick answer
Acceptance criteria for training data delivery are measurable pass/fail tests, agreed in the order or license schedule before any data moves, that decide whether a delivery is accepted, rejected or partly accepted. Each criterion needs a metric, the population or random sample it is measured on, a threshold, a named method or tool, who runs it, and the consequence of failure. A complete schedule covers file integrity, record counts, schema conformance, field completeness, duplicates, date coverage, label accuracy, residual personal data, and provenance and consent evidence.
By SourceX Editorial · Updated
The process around them (inspection windows, rejection notices, cure periods) is covered in dataset acceptance testing; the AI training data procurement hub shows where acceptance sits in the buying cycle.
What turns a quality promise into an acceptance criterion
A quality promise becomes an acceptance criterion when two people running the stated test on the same delivery would reach the same verdict. "Complete, high-quality ticket data" fails that test; "resolution_code non-missing in at least [y]% of closed tickets, checked on all records with the attached validator" passes. ISO/IEC 5259-2 sets out data quality measures [1], and one annotation vendor's guidance on dataset service terms likewise says to state how each term is measured and what remedy applies [2].
Write every criterion with seven parts:
- Unit: file, record, field value, label or document; per-label and per-record accuracy differ.
- Denominator: the exact population ("closed tickets created 2022-01-01 to 2024-12-31"), never "the dataset".
- Coverage: every record, or a random sample of stated size and selection method.
- Threshold: at least, at most or exactly, per stratum or overall.
- Method: validator, schema version, normalization rules or annotation guideline version.
- Who and when: buyer, supplier with a buyer witness, or an independent auditor, within a fixed window.
- Consequence: rejection, credit or replacement records; remedies for a defective data delivery compares them.
The guide to data supplier SLAs lists the surrounding service terms.
Gate checks that run on every file and record
Mechanical properties should be checked on 100% of the delivery, because software can test every record cheaply and sampling would only estimate what you can measure exactly. Treat them as gates: a failure stops further testing until re-delivery.
| Criterion | How it is measured | Failure it catches |
|---|---|---|
| File integrity | Every manifest file present with matching size and SHA-256 hash (manifest and checksum verification) | Missing shards, truncated or altered files |
| Parseability | JSON Lines: UTF-8 without a byte order mark, one valid JSON value per line, no blank lines [3]. Parquet: PAR1 magic number at start and end, readable footer metadata [4] | Byte order marks, partial lines, truncated files |
| Schema conformance | Every record validates against the agreed JSON Schema; as of October 2026 the current release is 2020-12, whose Validation specification defines keywords such as type, required and enum [5] | Renamed fields, numbers exported as strings, unlisted status values |
| Record counts | Exact count per file and per stratum (month, source system) against the license's record definition | Silent export filters, a missing month |
| Key integrity | Unique primary keys; every child record joins a parent | Duplicate IDs from merged exports, orphaned messages |
| Field completeness | Non-missing share per required field, with "missing" covering null, empty string and placeholders such as "N/A" or default dates | Empty fields, defaults masking gaps |
| Date coverage | Earliest and latest timestamps plus a maximum gap, such as no empty calendar month | A six-month hole inside the headline range |
| Duplicates | Exact matches by hash of normalized content; near-duplicates by a stated method and threshold | Re-exported batches, templated records |
Duplicate rates depend on the method. Lee et al. found one sentence repeated more than 60,000 times in C4, used both exact-substring and MinHash-based near-duplicate matching, and showed that training on deduplicated data cut verbatim emission of memorized text about tenfold [6]. Write the normalization, similarity measure and threshold into the criterion.
Attach the schema itself to the order.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "support_ticket_v3",
"type": "object",
"required": ["ticket_id", "created_at", "channel", "status", "messages"],
"properties": {
"ticket_id": {"type": "string", "pattern": "^T-[0-9]{8}$"},
"created_at": {"type": "string", "format": "date-time"},
"channel": {"enum": ["email", "chat", "phone_transcript", "web_form"]},
"status": {"enum": ["open", "pending", "solved", "closed"]},
"resolution_code": {"type": ["string", "null"]},
"messages": {
"type": "array",
"minItems": 1,
"items": {
"type": "object",
"required": ["author_role", "sent_at", "body"],
"properties": {
"author_role": {"enum": ["customer", "agent", "system"]},
"sent_at": {"type": "string", "format": "date-time"},
"body": {"type": "string", "minLength": 1}
}
}
}
}
}
Name the validator and its version, and state whether format keywords such as date-time fail a record or are only reported.
Sampled checks: label accuracy, field accuracy and residual personal data
Anything that needs human judgment, such as whether a label, transcription or redaction is correct, is measured on a random sample, so the criterion must state how many items are reviewed, how they are drawn, who adjudicates and how many errors the sample may contain, not only a target rate.
In manufacturing acceptance sampling, a single sampling plan is a pair (n, c): review n items and reject the lot if more than c are defective [7]. The NIST/SEMATECH handbook frames plans around the acceptable quality level (AQL), the producer's baseline that should pass with high probability, and the lot tolerance percent defective (LTPD), poor quality the consumer wants accepted only with very low probability [7][8]. ANSI/ASQ Z1.4 indexes plans by lot size, inspection level and AQL, with switching rules for a continuing stream of lots [9], which suits refresh deliveries. For a one-off purchase the buyer's main concern is the LTPD; acceptance sampling for dataset deliveries covers adapting these plans.
Define the defect before the number: a wrong label, an item with any wrong label, or a record with any wrong field, judged against a named guideline version by a named adjudicator.
Worked example: sizing a label-accuracy sample
Suppose the ML lead wants a delivery with 2% or more defective items accepted at most 5% of the time (LTPD 2%, consumer's risk 5%). These are example inputs, not recommendations. Binomial calculations for a delivery much larger than the sample give:
| Plan (n, c) | Accepts a 2% defective delivery | Accepts a 1% delivery | Accepts a 0.5% delivery |
|---|---|---|---|
| 149 items, 0 errors allowed | 4.9% | 22.4% | 47.4% |
| 313 items, 2 errors allowed | 5.0% | 39.3% | 79.3% |
| 456 items, 4 errors allowed | 4.9% | 52.0% | 91.9% |
All three plans protect the buyer equally at 2%. The zero-error plan, however, rejects a 0.5%-defect delivery more often than it accepts it, which invites disputes. Reviewing about three times as many items buys a plan a competent supplier can expect to pass. Record n, c and the LTPD in the schedule so either side can recompute them.
Three practices keep sampled criteria fair:
- Draw the sample yourself or through an independent party, from the full delivery, with a recorded random seed and strata where slices matter. Pre-payment acceptance audits are sold as an independent service [10].
- Count only human-confirmed defects. Northcutt et al. estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets, yet human reviewers confirmed only about 51% of the candidates their algorithm flagged [11].
- Do not read agreement as accuracy. Krippendorff's alpha on a double-annotated subset (1 is perfect reliability, 0 is chance-level agreement) [12] shows annotators were consistent; only an adjudicated sample shows they were correct.
Residual personal data works the same way. The Presidio project states that, because it relies on trained models, there is no guarantee it will find all sensitive information [13]. For HIPAA-covered data, OCR guidance recognizes Safe Harbor (removing 18 listed identifiers, with no actual knowledge that the rest could identify someone) and Expert Determination, and sets no numerical threshold for "very small" risk [14]. Require the method record or expert report for this exact delivery plus a sampled review; residual PII audit sampling sets out plans.
Provenance, consent and license evidence as acceptance conditions
A delivery whose records pass every data test should still fail if its records cannot be traced to a source the license covers. These criteria are checked against documents, so name the documents and how they join to the data.
- Source register coverage: every record's
source_idresolves to a register entry giving system of origin, collection period and permission basis. Threshold: all records. - Consent and exclusions: where the license relies on consent or a contractual permission, every record maps to that basis, and a full join against the opt-out and deletion list returns zero matches.
- License evidence, not labels: the Data Provenance Initiative's audit of more than 1,800 text datasets found license omission above 70% and error rates above 50% on popular hosting sites [15]. Require each source's permission documents.
- Documentation: a datasheet covering motivation, composition, collection process and recommended uses [16], or Croissant metadata, a schema.org-based JSON-LD vocabulary describing files and record-level structure [17] that a script can check against the delivery.
One dataset licensor's published process, for example, includes a final QA step covering consent records and metadata accuracy [18].
With SourceX, each dataset goes through rights review and is delivered under a license defining the included records, permitted uses, term and delivery method, and diligence materials on source, rights, preparation and allowed use are prepared per dataset. Personal details are removed or replaced before delivery and a sample is checked after processing, but no de-identification method is perfect, so keep a residual-PII criterion. See how SourceX works with data buyers.
Criteria that tighten for pre-training, fine-tuning, evaluation and RAG
The same records can pass for one use and fail for another, so tie each threshold to the use the license grants. For high-risk systems, EU AI Act Article 10(3) requires training, validation and testing data sets to be relevant, sufficiently representative and, to the best extent possible, free of errors and complete in view of the intended purpose [19]. As of October 2026, Regulation (EU) 2026/1744 has amended the AI Act; secondary sources report that Annex III high-risk obligations now apply from 2 December 2027 while the regulation also amends Article 10 [20].
| Use | Criteria to add or tighten | Why it matters |
|---|---|---|
| Pre-training | Near-duplicate rate with stated method; encoding and language checks; token counts from your named tokenizer; boilerplate share | Counts compare only when measured the same way; deduplication reduced memorized output [6] |
| Supervised fine-tuning | Turn structure (roles, order, no empty turns); adjudicated response accuracy; guideline version per batch | Malformed turns become training targets |
| Evaluation | Adjudicated gold labels; overlap checks against public benchmarks and your training data; minimum items per slice | Test-set label errors can change model rankings [11]; see accepting a delivered evaluation set |
| Retrieval (RAG) | Untruncated documents; headings and tables preserved; document ID, version and effective date | Without versions and dates, stale passages look current |
Illustrative acceptance schedule for a licensed support-ticket dataset
An acceptance schedule is the contract annex that lists each criterion with its population, method, threshold and consequence. This one fits a license of support-ticket histories for fine-tuning and evaluation; bracketed values are placeholders.
Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.
| # | Criterion | Measured on | Method | Threshold | On failure |
|---|---|---|---|---|---|
| 1 | Manifest integrity | All files | SHA-256 per file against the manifest | 100% match | Re-deliver affected files |
| 2 | Schema conformance | All records | JSON Schema support_ticket_v3, named validator and version | 100% valid | Reject lot |
| 3 | Record count | All records, by month | Closed tickets in the licensed date range | [N] ± [x]%; no month below [m] | Top-up delivery or pro-rata credit |
| 4 | Key integrity | All records | Unique ticket_id; every message joins a ticket | No duplicates or orphans | Reject lot |
| 5 | Field completeness | All closed tickets | Non-missing resolution_code, placeholders counted as missing | At least [y]% | Replacement records |
| 6 | Near-duplicates | All tickets | Stated normalization, similarity method and threshold | At most [d]% | Supplier removes and replaces |
| 7 | Intent label accuracy | Random sample, recorded seed, stratified by channel | Adjudicated review against guideline version [v] | (n, c) per stratum, from the agreed LTPD | Re-label the lot |
| 8 | Residual personal data | Random sample plus full-text scan | Human review of sampled and flagged tickets | No direct identifiers in the sample | Reject lot; re-run de-identification |
| 9 | Provenance | All records | source_id join to source register; exclusion-list join | All resolve; no excluded subjects | Reject lot |
| 10 | Documentation | Delivery package | Datasheet, preparation record, schema and manifest present and consistent | All present | Window starts when complete |
Format, transfer channel and security terms belong in the technical delivery specification, which the schedule should reference rather than repeat. If payment is staged, map gate checks and sampled checks to separate milestone payments tied to acceptance.
Drafting mistakes that make acceptance criteria unenforceable
Acceptance disputes often trace back to a target without a denominator, a method or a sample design. Watch for:
- Averages hiding a failed slice. Set per-stratum thresholds wherever a slice matters to the model.
- Criteria proven only on the pre-contract sample. Results carry over only if the full set is drawn and prepared the same way; request the sample on that basis.
- Deemed acceptance by silence. Defects often surface later, during training.
- Unversioned criteria. Schema or guideline changes between refreshes need a matching schedule version; recurring service levels belong in SLAs for recurring data deliveries.
Writing acceptance terms for a data purchase?
If you want acceptance tests settled before anything is delivered, describe the records you need, the uses you want licensed and the thresholds that matter to your model. SourceX looks for US companies that hold that data, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is delivered until an agreement is executed and the supplier approves the terms. Describe your data and acceptance requirements.
Sources
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- Digital Divide Data, "Dataset service terms: measurement and remedies" (vendor blog; descriptive title; market practice only). https://www.digitaldividedata.com/?p=24243
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- The Apache Software Foundation, "File Format" (Apache Parquet documentation). https://parquet.apache.org/docs/file-format/
- JSON Schema, "Specification". https://json-schema.org/specification
- Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021; ACL 2022). https://arxiv.org/abs/2107.06499v1
- NIST/SEMATECH, "How do you choose a single sampling plan?" (e-Handbook of Statistical Methods, section 6.2.3). https://itl.nist.gov/div898/handbook/pmc/section2/pmc23.htm
- NIST/SEMATECH, "Choosing a single sampling plan" (e-Handbook of Statistical Methods, section 6.2.3.1). https://itl.nist.gov/div898/handbook/pmc/section2/pmc231.htm
- ASQ Quality Press, "ASQ/ANSI Z1.4:2003 (R2018): Sampling Procedures and Tables for Inspection by Attributes". https://asq.org/quality-press/display-item?item=T1164
- Contra, "Robot dataset acceptance audit: a verdict before you pay" (freelance service listing; market practice only). https://contra.com/s/ee6vvEQU-robot-dataset-acceptance-audit-a-verdict-before-you-pay
- Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Krippendorff, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
- Microsoft, "Presidio - Data Protection API" (project documentation on pkg.go.dev). https://pkg.go.dev/github.com/microsoft/presidio
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version in Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- Gebru et al., "Datasheets for Datasets" (2018; Communications of the ACM 2021). https://arxiv.org/pdf/1803.09010
- Akhtar et al., "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- Pocstock, "The dataset licensing process from inquiry to delivery" (help center; market practice only). https://support.pocstock.com/en/articles/14772883-the-dataset-licensing-process-from-inquiry-to-delivery
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- European Parliament and Council of the European Union, "Regulation (EU) 2026/1744 amending Regulations (EU) 2024/1689, (EU) 2018/1139 and (EU) 2023/1230 (Digital Omnibus on AI)" (Official Journal, 24 July 2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.