Skip to content

Procurement, samples and ongoing supply

Acceptance criteria for training data delivery: thresholds to write into the order

Quick answer

Acceptance criteria for training data delivery are measurable pass/fail tests, agreed in the order or license schedule before any data moves, that decide whether a delivery is accepted, rejected or partly accepted. Each criterion needs a metric, the population or random sample it is measured on, a threshold, a named method or tool, who runs it, and the consequence of failure. A complete schedule covers file integrity, record counts, schema conformance, field completeness, duplicates, date coverage, label accuracy, residual personal data, and provenance and consent evidence.

By SourceX Editorial · Updated

The process around them (inspection windows, rejection notices, cure periods) is covered in dataset acceptance testing; the AI training data procurement hub shows where acceptance sits in the buying cycle.

What turns a quality promise into an acceptance criterion

A quality promise becomes an acceptance criterion when two people running the stated test on the same delivery would reach the same verdict. "Complete, high-quality ticket data" fails that test; "resolution_code non-missing in at least [y]% of closed tickets, checked on all records with the attached validator" passes. ISO/IEC 5259-2 sets out data quality measures [1], and one annotation vendor's guidance on dataset service terms likewise says to state how each term is measured and what remedy applies [2].

Write every criterion with seven parts:

  1. Unit: file, record, field value, label or document; per-label and per-record accuracy differ.
  2. Denominator: the exact population ("closed tickets created 2022-01-01 to 2024-12-31"), never "the dataset".
  3. Coverage: every record, or a random sample of stated size and selection method.
  4. Threshold: at least, at most or exactly, per stratum or overall.
  5. Method: validator, schema version, normalization rules or annotation guideline version.
  6. Who and when: buyer, supplier with a buyer witness, or an independent auditor, within a fixed window.
  7. Consequence: rejection, credit or replacement records; remedies for a defective data delivery compares them.

The guide to data supplier SLAs lists the surrounding service terms.

Gate checks that run on every file and record

Mechanical properties should be checked on 100% of the delivery, because software can test every record cheaply and sampling would only estimate what you can measure exactly. Treat them as gates: a failure stops further testing until re-delivery.

CriterionHow it is measuredFailure it catches
File integrityEvery manifest file present with matching size and SHA-256 hash (manifest and checksum verification)Missing shards, truncated or altered files
ParseabilityJSON Lines: UTF-8 without a byte order mark, one valid JSON value per line, no blank lines [3]. Parquet: PAR1 magic number at start and end, readable footer metadata [4]Byte order marks, partial lines, truncated files
Schema conformanceEvery record validates against the agreed JSON Schema; as of October 2026 the current release is 2020-12, whose Validation specification defines keywords such as type, required and enum [5]Renamed fields, numbers exported as strings, unlisted status values
Record countsExact count per file and per stratum (month, source system) against the license's record definitionSilent export filters, a missing month
Key integrityUnique primary keys; every child record joins a parentDuplicate IDs from merged exports, orphaned messages
Field completenessNon-missing share per required field, with "missing" covering null, empty string and placeholders such as "N/A" or default datesEmpty fields, defaults masking gaps
Date coverageEarliest and latest timestamps plus a maximum gap, such as no empty calendar monthA six-month hole inside the headline range
DuplicatesExact matches by hash of normalized content; near-duplicates by a stated method and thresholdRe-exported batches, templated records

Duplicate rates depend on the method. Lee et al. found one sentence repeated more than 60,000 times in C4, used both exact-substring and MinHash-based near-duplicate matching, and showed that training on deduplicated data cut verbatim emission of memorized text about tenfold [6]. Write the normalization, similarity measure and threshold into the criterion.

Attach the schema itself to the order.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "title": "support_ticket_v3",
  "type": "object",
  "required": ["ticket_id", "created_at", "channel", "status", "messages"],
  "properties": {
    "ticket_id": {"type": "string", "pattern": "^T-[0-9]{8}$"},
    "created_at": {"type": "string", "format": "date-time"},
    "channel": {"enum": ["email", "chat", "phone_transcript", "web_form"]},
    "status": {"enum": ["open", "pending", "solved", "closed"]},
    "resolution_code": {"type": ["string", "null"]},
    "messages": {
      "type": "array",
      "minItems": 1,
      "items": {
        "type": "object",
        "required": ["author_role", "sent_at", "body"],
        "properties": {
          "author_role": {"enum": ["customer", "agent", "system"]},
          "sent_at": {"type": "string", "format": "date-time"},
          "body": {"type": "string", "minLength": 1}
        }
      }
    }
  }
}

Name the validator and its version, and state whether format keywords such as date-time fail a record or are only reported.

Sampled checks: label accuracy, field accuracy and residual personal data

Anything that needs human judgment, such as whether a label, transcription or redaction is correct, is measured on a random sample, so the criterion must state how many items are reviewed, how they are drawn, who adjudicates and how many errors the sample may contain, not only a target rate.

In manufacturing acceptance sampling, a single sampling plan is a pair (n, c): review n items and reject the lot if more than c are defective [7]. The NIST/SEMATECH handbook frames plans around the acceptable quality level (AQL), the producer's baseline that should pass with high probability, and the lot tolerance percent defective (LTPD), poor quality the consumer wants accepted only with very low probability [7][8]. ANSI/ASQ Z1.4 indexes plans by lot size, inspection level and AQL, with switching rules for a continuing stream of lots [9], which suits refresh deliveries. For a one-off purchase the buyer's main concern is the LTPD; acceptance sampling for dataset deliveries covers adapting these plans.

Define the defect before the number: a wrong label, an item with any wrong label, or a record with any wrong field, judged against a named guideline version by a named adjudicator.

Worked example: sizing a label-accuracy sample

Suppose the ML lead wants a delivery with 2% or more defective items accepted at most 5% of the time (LTPD 2%, consumer's risk 5%). These are example inputs, not recommendations. Binomial calculations for a delivery much larger than the sample give:

Plan (n, c)Accepts a 2% defective deliveryAccepts a 1% deliveryAccepts a 0.5% delivery
149 items, 0 errors allowed4.9%22.4%47.4%
313 items, 2 errors allowed5.0%39.3%79.3%
456 items, 4 errors allowed4.9%52.0%91.9%

All three plans protect the buyer equally at 2%. The zero-error plan, however, rejects a 0.5%-defect delivery more often than it accepts it, which invites disputes. Reviewing about three times as many items buys a plan a competent supplier can expect to pass. Record n, c and the LTPD in the schedule so either side can recompute them.

Three practices keep sampled criteria fair:

  • Draw the sample yourself or through an independent party, from the full delivery, with a recorded random seed and strata where slices matter. Pre-payment acceptance audits are sold as an independent service [10].
  • Count only human-confirmed defects. Northcutt et al. estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets, yet human reviewers confirmed only about 51% of the candidates their algorithm flagged [11].
  • Do not read agreement as accuracy. Krippendorff's alpha on a double-annotated subset (1 is perfect reliability, 0 is chance-level agreement) [12] shows annotators were consistent; only an adjudicated sample shows they were correct.

Residual personal data works the same way. The Presidio project states that, because it relies on trained models, there is no guarantee it will find all sensitive information [13]. For HIPAA-covered data, OCR guidance recognizes Safe Harbor (removing 18 listed identifiers, with no actual knowledge that the rest could identify someone) and Expert Determination, and sets no numerical threshold for "very small" risk [14]. Require the method record or expert report for this exact delivery plus a sampled review; residual PII audit sampling sets out plans.

A delivery whose records pass every data test should still fail if its records cannot be traced to a source the license covers. These criteria are checked against documents, so name the documents and how they join to the data.

  • Source register coverage: every record's source_id resolves to a register entry giving system of origin, collection period and permission basis. Threshold: all records.
  • Consent and exclusions: where the license relies on consent or a contractual permission, every record maps to that basis, and a full join against the opt-out and deletion list returns zero matches.
  • License evidence, not labels: the Data Provenance Initiative's audit of more than 1,800 text datasets found license omission above 70% and error rates above 50% on popular hosting sites [15]. Require each source's permission documents.
  • Documentation: a datasheet covering motivation, composition, collection process and recommended uses [16], or Croissant metadata, a schema.org-based JSON-LD vocabulary describing files and record-level structure [17] that a script can check against the delivery.

One dataset licensor's published process, for example, includes a final QA step covering consent records and metadata accuracy [18].

With SourceX, each dataset goes through rights review and is delivered under a license defining the included records, permitted uses, term and delivery method, and diligence materials on source, rights, preparation and allowed use are prepared per dataset. Personal details are removed or replaced before delivery and a sample is checked after processing, but no de-identification method is perfect, so keep a residual-PII criterion. See how SourceX works with data buyers.

Criteria that tighten for pre-training, fine-tuning, evaluation and RAG

The same records can pass for one use and fail for another, so tie each threshold to the use the license grants. For high-risk systems, EU AI Act Article 10(3) requires training, validation and testing data sets to be relevant, sufficiently representative and, to the best extent possible, free of errors and complete in view of the intended purpose [19]. As of October 2026, Regulation (EU) 2026/1744 has amended the AI Act; secondary sources report that Annex III high-risk obligations now apply from 2 December 2027 while the regulation also amends Article 10 [20].

UseCriteria to add or tightenWhy it matters
Pre-trainingNear-duplicate rate with stated method; encoding and language checks; token counts from your named tokenizer; boilerplate shareCounts compare only when measured the same way; deduplication reduced memorized output [6]
Supervised fine-tuningTurn structure (roles, order, no empty turns); adjudicated response accuracy; guideline version per batchMalformed turns become training targets
EvaluationAdjudicated gold labels; overlap checks against public benchmarks and your training data; minimum items per sliceTest-set label errors can change model rankings [11]; see accepting a delivered evaluation set
Retrieval (RAG)Untruncated documents; headings and tables preserved; document ID, version and effective dateWithout versions and dates, stale passages look current

Illustrative acceptance schedule for a licensed support-ticket dataset

An acceptance schedule is the contract annex that lists each criterion with its population, method, threshold and consequence. This one fits a license of support-ticket histories for fine-tuning and evaluation; bracketed values are placeholders.

Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.

#CriterionMeasured onMethodThresholdOn failure
1Manifest integrityAll filesSHA-256 per file against the manifest100% matchRe-deliver affected files
2Schema conformanceAll recordsJSON Schema support_ticket_v3, named validator and version100% validReject lot
3Record countAll records, by monthClosed tickets in the licensed date range[N] ± [x]%; no month below [m]Top-up delivery or pro-rata credit
4Key integrityAll recordsUnique ticket_id; every message joins a ticketNo duplicates or orphansReject lot
5Field completenessAll closed ticketsNon-missing resolution_code, placeholders counted as missingAt least [y]%Replacement records
6Near-duplicatesAll ticketsStated normalization, similarity method and thresholdAt most [d]%Supplier removes and replaces
7Intent label accuracyRandom sample, recorded seed, stratified by channelAdjudicated review against guideline version [v](n, c) per stratum, from the agreed LTPDRe-label the lot
8Residual personal dataRandom sample plus full-text scanHuman review of sampled and flagged ticketsNo direct identifiers in the sampleReject lot; re-run de-identification
9ProvenanceAll recordssource_id join to source register; exclusion-list joinAll resolve; no excluded subjectsReject lot
10DocumentationDelivery packageDatasheet, preparation record, schema and manifest present and consistentAll presentWindow starts when complete

Format, transfer channel and security terms belong in the technical delivery specification, which the schedule should reference rather than repeat. If payment is staged, map gate checks and sampled checks to separate milestone payments tied to acceptance.

Drafting mistakes that make acceptance criteria unenforceable

Acceptance disputes often trace back to a target without a denominator, a method or a sample design. Watch for:

  • Averages hiding a failed slice. Set per-stratum thresholds wherever a slice matters to the model.
  • Criteria proven only on the pre-contract sample. Results carry over only if the full set is drawn and prepared the same way; request the sample on that basis.
  • Deemed acceptance by silence. Defects often surface later, during training.
  • Unversioned criteria. Schema or guideline changes between refreshes need a matching schedule version; recurring service levels belong in SLAs for recurring data deliveries.

Writing acceptance terms for a data purchase?

If you want acceptance tests settled before anything is delivered, describe the records you need, the uses you want licensed and the thresholds that matter to your model. SourceX looks for US companies that hold that data, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is delivered until an agreement is executed and the supplier approves the terms. Describe your data and acceptance requirements.

Sources

  1. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
  2. Digital Divide Data, "Dataset service terms: measurement and remedies" (vendor blog; descriptive title; market practice only). https://www.digitaldividedata.com/?p=24243
  3. jsonlines.org, "JSON Lines". https://jsonlines.org/
  4. The Apache Software Foundation, "File Format" (Apache Parquet documentation). https://parquet.apache.org/docs/file-format/
  5. JSON Schema, "Specification". https://json-schema.org/specification
  6. Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021; ACL 2022). https://arxiv.org/abs/2107.06499v1
  7. NIST/SEMATECH, "How do you choose a single sampling plan?" (e-Handbook of Statistical Methods, section 6.2.3). https://itl.nist.gov/div898/handbook/pmc/section2/pmc23.htm
  8. NIST/SEMATECH, "Choosing a single sampling plan" (e-Handbook of Statistical Methods, section 6.2.3.1). https://itl.nist.gov/div898/handbook/pmc/section2/pmc231.htm
  9. ASQ Quality Press, "ASQ/ANSI Z1.4:2003 (R2018): Sampling Procedures and Tables for Inspection by Attributes". https://asq.org/quality-press/display-item?item=T1164
  10. Contra, "Robot dataset acceptance audit: a verdict before you pay" (freelance service listing; market practice only). https://contra.com/s/ee6vvEQU-robot-dataset-acceptance-audit-a-verdict-before-you-pay
  11. Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  12. Krippendorff, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  13. Microsoft, "Presidio - Data Protection API" (project documentation on pkg.go.dev). https://pkg.go.dev/github.com/microsoft/presidio
  14. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  15. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version in Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
  16. Gebru et al., "Datasheets for Datasets" (2018; Communications of the ACM 2021). https://arxiv.org/pdf/1803.09010
  17. Akhtar et al., "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  18. Pocstock, "The dataset licensing process from inquiry to delivery" (help center; market practice only). https://support.pocstock.com/en/articles/14772883-the-dataset-licensing-process-from-inquiry-to-delivery
  19. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  20. European Parliament and Council of the European Union, "Regulation (EU) 2026/1744 amending Regulations (EU) 2024/1689, (EU) 2018/1139 and (EU) 2023/1230 (Digital Omnibus on AI)" (Official Journal, 24 July 2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data