Skip to content

Procurement, samples and ongoing supply

How to request a training data sample from a supplier

Quick answer

To request a data sample from a vendor that you can actually test, send a written specification rather than asking for "a sample." Name the tests the sample must support and the size each needs, and say how records are selected (random or stratified, with the seed and query). Require the same de-identification, filtering, schema and file format as the full delivery. Ask for a data dictionary and full-population statistics with it, and agree evaluation terms that permit your tests.

By SourceX Editorial · Updated

Match the sample stage to the decision it supports

A sample request starts from the decision the sample informs, because a schema preview, a look sample, an evaluation sample and a trial slice prove different things. A purchase can pass through several of these stages between RFI and contract in the AI training data procurement lifecycle.

StageWhat you receiveWhat it can showWhat it cannot show
Schema previewData dictionary, record counts, a few dummy rowsWhether fields, joins and time span fitReal values, quality, privacy residue
Look sampleTens of real recordsFormats, free-text style, how identifiers were replacedAny rate; whether the sample is typical
Evaluation sampleHundreds to thousands of records drawn by a stated methodDefect rates, label quality, distribution match, small fine-tune or retrieval effectsLong-tail coverage, unless you specified strata
Trial or pilot sliceA bounded slice of the product, such as one region or time windowEnd-to-end ingestion and model lift in your own pipelineBehavior outside the trial window

Conventions differ by market. The FISD Alternative Data Council's illustrative due diligence questionnaire, written for data sold to investment managers, asks providers for a data dictionary and a sample of fewer than 100 rows from data more than three months old [1], which is a look sample: enough to read the schema, too small to measure anything. One licensor's documented process includes a sample stage in which the buyer evaluates the sample and any iterations happen before final delivery [2].

If the holder will not release raw records, the options narrow to secure viewing, redacted records or synthetic stand-ins; see what restricted sample routes can and cannot prove.

How big a data sample should be for each test

There is no single answer to how big a data sample should be: size it from the test it feeds. A schema check needs every record type, a defect-rate check enough random records to rule out the rate you care about, and a model test enough examples to move a metric your evaluation set can resolve.

Test you plan to runWhat sets the sizePractical floor
Parsing and schema conformanceCoverage of every record type and optional fieldEach variant several times
Defect rate (residual PII, duplicates)The highest rate you would acceptAbout 300 random records to reject a 1% rate when none are found (worked example below)
Label accuracyExpected error rate and needed precisionHundreds of expert-reviewed items; test sets of 10 widely used datasets averaged an estimated 3.3% or more label errors [4]
Annotator agreementItems labeled independently by two or more annotatorsA multiply-annotated subset large enough to compute Krippendorff's alpha per label [5]
Fine-tuning ablationWhether an effect shows on your evaluation setAbout a thousand examples can change behavior: LIMA fine-tuned a 65B-parameter model on 1,000 curated pairs [6]
Sample reused as evaluation dataWidth of the confidence intervalAt least a few hundred items; below that, standard CLT-based error bars tend to be too narrow [7]
Retrieval (RAG) testCorpus density and distractors around your queriesA complete slice of documents, not scattered chunks

Your evaluation set limits what an ablation can show: one practitioner guide calculates that detecting a lift from 82% to 85% accuracy at 80% power and 5% significance needs about 2,400 labeled examples per model variant [8]. Measurement methods are in estimating a dataset's value before you buy.

For repeated deliveries, ANSI/ASQ Z1.4 tables give sample sizes and accept/reject numbers by lot size, inspection level and acceptable quality level (AQL) for attribute inspection [3]. The standard was written for product lots, so define how a delivery batch and a defective record map onto it in your acceptance criteria for licensed training data.

Worked example: sizing one request for three tests

Illustrative example: invented to show structure; it does not describe an available dataset.

A team plans a residual-PII check, a supervised fine-tuning (SFT) ablation and a retrieval test on a support-ticket dataset. The supplier says fewer than 1% of records retain an identifier.

  1. PII check. If the true rate were 1%, the chance that n random records contain no identifier is 0.99^n: about 37% at n = 100, about 5% at n = 300. A clean 300-record review therefore rejects "1% or worse" with roughly 95% confidence. Bounds for other rates are in sample sizes for estimating a dataset's error rate.
  2. SFT ablation. Ask for 2,000 records by simple random draw and review 300 of them for step 1, so one draw serves both tests.
  3. Tail coverage. Add 400 records oversampled from three rare strata (escalations, refunds, non-English tickets), each tagged with its sampling weight.
  4. Retrieval test. Request every ticket and linked knowledge-base article for one product line and quarter, because a retrieval test needs the documents that compete with the right answer.

Specify how records are selected, not only how many

The selection method decides whether results transfer to the full dataset, so the request names it: a simple random draw with a recorded seed, or a stratified draw across strata you choose. A supplier-curated "best of" set answers none of your questions. Ask for the selection query or script, the seed and the population count behind each stratum.

Pick strata that drive model behavior in your use case: source system, year or quarter, record type, language, customer segment, outcome (resolved, escalated, refunded) and length bands for free text. Then choose an allocation. Research on data representativity separates two goals: a sample that mirrors the population, which supports average-case estimates, and one that covers the input space, which helps robustness to distribution shift and accuracy on smaller subgroups [9]. A proportional random core plus a weighted oversample of the tail can serve both goals.

Ask for full-population statistics in the same request: counts per stratum, field fill rates, date histograms, text-length percentiles, label distributions and records removed by each filter. The comparison method is in checking whether a vendor's sample is representative, tracing sampled records to consent evidence is in testing a supplier's provenance claims on a sample, and the supplier's view is in the supplier-side guide to selecting a representative sample.

Require the sample to be prepared exactly like the full delivery

A sample predicts the full dataset only if it passes through the same preparation pipeline, at the same version, as the records you will license. A hand-cleaned or separately exported sample hides the defects you are paying to find. Spell out parity on five points:

  • De-identification. The same method, entity types, replacement scheme (consistent surrogate values or generic masks) and tool version. Automated detection misses things: the open-source Presidio project cautions that because it uses trained models there is no guarantee it will find all sensitive information [10]. For protected health information covered by HIPAA, name the method: Safe Harbor removal of 18 identifier types, or Expert Determination [11].
  • Filters and deduplication. The same language, date and quality filters, applied before sampling. Near-duplicates are common in web-scale text corpora; one study found a single sentence repeated more than 60,000 times in the C4 corpus [12]. A sample drawn before deduplication misstates the duplicate rate you will receive.
  • Schema and IDs. The schema version, field names and stable record IDs that will appear in the full delivery, so every sample finding maps to a delivered record.
  • File format and writer settings. The same container, compression and writer. Parquet implementations do not all support the same features of the format [13], so a sample written by one library can read cleanly while the full delivery fails in yours.
  • Manifest. File list, record counts and checksums in the form you expect at delivery. An illustrative package summary and field guide shows the level of detail to ask for.

Documents to request with the sample

Ask for the documents that let you interpret the records: a data dictionary, a selection log, a preparation log, population statistics and a datasheet. A supplier who knows its data can produce them, so a gap is itself a finding.

  • Data dictionary: field names, types, units, allowed values, null semantics and code lists. The FISD template asks for one alongside its sample [1].
  • Selection log: query, seed, strata, weights and population counts.
  • Preparation log: de-identification method and version, filters, deduplication method and threshold, records removed per step.
  • Population statistics for the full candidate dataset, as listed above.
  • Datasheet answers on motivation, composition, collection process and recommended uses, following Datasheets for Datasets [14].
  • Machine-readable metadata where available, such as Croissant, a schema.org-based JSON-LD vocabulary covering dataset, file and record structure [15]. As of October 2026, NeurIPS requires responsible-AI metadata based on Croissant RAI fields for its Evaluations and Datasets Track [16], a sign of where documentation norms are heading.
  • For labeled data: annotation guidelines, the multiply-annotated subset and agreement statistics.
  • Known issues: drifting fields, date gaps, mid-period system changes.

The wider pre-signature evidence pack, including profiling reports and provenance summaries, is covered in evidence to request from a data vendor before you sign. If you source through SourceX, you can put this list straight into a data request on the SourceX buyer page. SourceX looks for US businesses that hold the data you describe, and prepares diligence materials on source, rights, preparation and allowed use per dataset for your review.

Sample terms that permit the tests you planned

Before the sample ships, confirm in writing that the evaluation terms allow every test in your request, including any fine-tuning run, and state what happens afterward to the records, checkpoints and results; evaluation grants are often narrower than buyers assume.

NVIDIA's sample data license for evaluation (2026.01.19 version) grants a limited, non-exclusive, revocable, non-transferable, non-sublicensable right solely to evaluate and test NVIDIA technologies, and bars other uses and redistribution [17]. A grant worded that way would not cover training your own model on the sample. One data vendor's template combines a trial data license and a mutual NDA in one document [18]. Check five points:

  1. Permitted tests. Name each test, plus the internal teams and cloud environments involved.
  2. Results. You keep metrics, error analyses and evaluation reports, whether or not you buy.
  3. Retention. A deletion date for the sample and any checkpoints trained on it, with written certification.
  4. Overlap with the full delivery. Sample record IDs flagged in the full delivery, so records you used for evaluation stay out of training; see contamination checks for licensed evaluation data.
  5. Re-identification. As of October 2026, under the CCPA definition in California Civil Code section 1798.140(m), information counts as deidentified only if, among other conditions, the business contractually obligates any recipients to comply with all of the definition's provisions, which include not attempting to re-identify it [19]. A supplier relying on that status therefore has reason to include the clause in sample terms.

Negotiating these terms is covered in evaluation licenses and NDAs for dataset samples, and fees and conversion credits for larger trials in paid data pilot terms. Suppliers' own expectations are described in the supplier-side explanation of sharing samples.

Illustrative dataset sample request

Putting size, selection, preparation, documents and terms in one written request lets you compare suppliers' answers line by line.

Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.

sample_request:
  candidate_dataset: "B2B software support tickets with agent replies, 2021-2025"
  planned_tests: [pii_residue_review, sft_ablation, retrieval_test]
  selection:
    core:
      records: 2000
      method: simple_random
      seed: 20261009
    tail_oversample:
      records: 400
      strata: [escalated_to_engineering, refund_issued, non_english]
      report_sampling_weights: true
    retrieval_slice: "all tickets and linked KB articles, one product line, one quarter"
  preparation_parity:
    pipeline_version: same_as_full_delivery
    deidentification: "same method, entity list and surrogate scheme; method and version named"
    filters_and_dedup: "applied before sampling; thresholds stated"
    format: "JSONL, UTF-8, one ticket per line; same schema version and record IDs"
  documents:
    - data_dictionary
    - selection_log          # query, seed, strata, population counts
    - preparation_log        # steps, versions, records removed per step
    - population_statistics  # per stratum: counts, fill rates, date histogram, length percentiles
    - datasheet
    - manifest_with_checksums
  terms:
    permitted_tests: [profiling, manual_review, fine_tuning_ablation, retrieval_test]
    buyer_keeps_results: true
    delete_sample_and_checkpoints: "within 60 days of decision, certified"
    flag_sample_ids_in_full_delivery: true
    no_reidentification: true

Red flags when the sample arrives

Treat a sample as a preview of how the supplier will behave at delivery, so read these signs as findings to resolve before you sign:

  • No selection log, or "selected for quality" given as the method.
  • Every record from one year, customer or system when the dataset is described as multi-year or multi-source.
  • A different format, schema version or ID scheme from the one quoted for delivery.
  • Zero residual identifiers and no preparation log, which suggests hand-cleaning.
  • Population statistics that cannot be reconciled with the sample's stratum counts.
  • Evaluation terms that arrive after the data, or that forbid the tests you named.

Rates measured on a sample that passes become the baseline for delivery thresholds. Modality-specific tests are covered in evaluating a code dataset sample, evaluating an agent data sample and pilot-testing a content source for retrieval lift.

Describe the dataset and sample you need

If you need operational records from US businesses, describe the dataset and the sample specification you will test against on the SourceX buyer page. SourceX looks for US companies that hold the data you describe, checks the data and the supplier's licensing permissions, and manages the license and delivery; a request does not guarantee a matching dataset. Describe the data and sample you need.

Sources

  1. FISD Alternative Data Council, "Data Provider Due Diligence Questionnaire (DDQ), with generative AI questions" (2024). https://fisd.net/wp-content/uploads/2024/02/FISD-Alternative-Data-Council-Due-Diligence-Questionnaire-with-GenAI-Questions-022824.docx
  2. Pocstock, "The dataset licensing process from inquiry to delivery". https://support.pocstock.com/en/articles/14772883-the-dataset-licensing-process-from-inquiry-to-delivery
  3. ASQ / ANSI, "ASQ/ANSI Z1.4:2003 (R2018): Sampling Procedures and Tables for Inspection by Attributes" (2003, reaffirmed 2018). https://asq.org/quality-press/display-item?item=T1164
  4. Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  5. Klaus Krippendorff, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  6. Zhou et al., "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  7. ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  8. tianpan.co, "Statistical power for LLM evals" (2026). https://tianpan.co/blog/2026/04/15/statistical-power-llm-evals
  9. arXiv, "Data Representativity for Machine Learning and AI Systems" (2022). https://arxiv.org/pdf/2203.04706
  10. microsoft/presidio project, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  11. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  12. Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  13. The Apache Software Foundation, "Apache Parquet Documentation". https://parquet.apache.org/docs/file-format/implementationstatus/
  14. Gebru et al., "Datasheets for Datasets" (2018; CACM 2021). https://arxiv.org/pdf/1803.09010
  15. Akhtar et al., MLCommons, "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  16. NeurIPS blog, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
  17. NVIDIA, "NVIDIA Sample Data License for Evaluation (2026.01.19)" (2026). https://developer.download.nvidia.com/licenses/nvidia-sample-data-license-for-evaluation-2026.01.19.pdf
  18. New Constructs, "Trial Data License Agreement and Mutual Non-Disclosure Agreement (TDLA Mutual NDA General)" (2024). https://www.newconstructs.com/wp-content/uploads/2024/10/New-Constructs-TDLA-Mutual-NDA-General.pdf
  19. California Legislature, "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data