Skip to content

Procurement, samples and ongoing supply

Evidence to Request from a Data Vendor Before You Sign

Quick answer

Before signing, request seven artifacts: a representative sample under an evaluation license, a field-level data dictionary, a profiling report computed on the full delivery population, a provenance and rights summary, a written de-identification method with its validation results, labeling and QA records, and test results tied to a named dataset revision. Each should let your team check a claim without trusting the vendor. Collect them before signature, because after kickoff your leverage to demand them drops sharply [1].

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why an approval file needs artifacts, not assurances

An approval file should hold things a reviewer can re-check, because a vendor statement on its own proves nothing [3]. "High quality," "diverse" and "fine-tuning ready" are opinions; "every line parses as JSONL against this schema" and "exact-duplicate rate below 1% on the SHA-256 of the normalized text field" are facts that can be tested [2]. The useful rule is to convert each marketing claim into a measurable statement, then ask which artifact would prove it.

This page covers that evidence pack. If you are still shortlisting, start with the questions to ask a training data vendor; to weight vendors against each other, use the data vendor evaluation scorecard. The broader procurement guide to AI training data shows where this step sits between RFP and onboarding.

The seven artifacts and what each must prove

Each artifact answers one question a reviewer will ask, and each has a typical way of being incomplete. Request them by name in writing, and record in the approval file which revision of each you received. The AI training data due diligence checklist covers the wider review, and dataset cards for licensed enterprise data shows a documentation format.

  1. Representative sample. Proves the format, content and label quality match the description. Ask how it was drawn (random seed, stratification keys) so you can judge whether it reflects the full set; checking whether a sample is representative covers the tests.
  2. Data dictionary. Proves you know what every field means: name, type, allowed values, null semantics, units, time zone, and which fields are derived. Industry DDQs such as the FISD Alternative Data Council template treat the dictionary and a sample as baseline diligence items [4].
  3. Profiling report on the full population. Proves the sample's statistics hold at scale: row counts, null rates per field, cardinality, value distributions, date ranges, language mix, and duplicate rates. It must be generated on the deliverable, not the sample.
  4. Provenance and rights summary. Proves where records came from, under what terms, and what consents or contracts cover them. A Data Card style summary of upstream sources, collection method, annotation method and intended use is a reasonable template [5].
  5. De-identification method and validation. Proves personal data handling is documented and tested, not just asserted. For US health data, the method must name HIPAA Safe Harbor or Expert Determination [7].
  6. Labeling and QA records. Proves annotations meet a stated standard: guidelines version, annotator qualification, inter-annotator agreement, adjudication rate. ISO/IEC 5259-4 frames this as part of a data quality process for training and evaluation data [6].
  7. Revision-tied test results. Proves the checks above apply to the exact files you will receive, identified by version tag or manifest hash.

Turning vendor claims into checkable statements

The fastest way to build the evidence request is to rewrite each claim in the vendor's deck as a metric, a threshold and a method. Claims that cannot be rewritten this way go in the file as unverified opinion, not as evidence [2].

Vendor claimCheckable restatementArtifact that proves itYour check
"Clean, deduplicated"Exact-duplicate rate below 1% and near-duplicate rate (MinHash, Jaccard 0.8) below 3% on the primary text fieldProfiling report with method and thresholdsRe-run hashing on the sample
"Valid, ready to load"100% of records parse as JSONL or Parquet against the published JSON Schema or Arrow schemaData dictionary plus schema fileValidate the sample with your loader
"Expert-labeled"Labels by annotators with named qualifications; Cohen's kappa or Krippendorff's alpha reported per labelQA records and guideline versionBlind re-label 100 sample items
"PII removed"Named fields masked or tokenized; free text scanned with a stated detector; residual hit rate on a reviewed sampleDe-identification method and validation memoRun your own PII scan on the sample
"Representative of production"Distribution by stated strata (region, product line, year) matches the population within stated toleranceProfiling report and sampling noteCompare sample and full-set marginals
"Good for fine-tuning"Not checkable as statedNoneRun your own bounded test

Which evaluation the vendor's test results actually cover

Vendor test results usually describe one of three different things, and the approval file should say which [2]. The first is a revision-level test: checks that ran on a specific dataset version, such as schema validation, duplicate scans and PII scans. The second is the producer's correction process: how errors found after delivery are reported, fixed and re-issued. The third is fitness for your model and task, which only you can measure.

Ask for the first two in writing and keep the third in your own hands. A bounded proof on your own evaluation set, such as a small fine-tune scored on a held-out set you built before seeing the data, is stronger evidence than any vendor benchmark [3]. For fine-tuning data specifically, evaluating a fine-tuning dataset before you buy it gives a test design, and estimating a dataset's value to your model covers how to turn results into a price ceiling.

Provenance, rights and de-identification evidence

Rights evidence is the artifact most often replaced by a sentence like "all data is properly licensed," so ask for the chain behind it. A usable provenance summary names the system each record came from (for example a Zendesk ticket export, a Salesforce opportunity history or a Jira project), the collection period, the legal basis or contract under which the supplier holds it, and whether end users or employees consented to secondary use. If you develop or substantially modify a generative AI system made available to Californians, AB 2013 has required you to post high-level training data documentation since January 1, 2026 (status as of October 2026), so your own disclosure depends on what the vendor can document [8].

For personal data, request the de-identification method as a document, not a checkbox. It should list the direct identifiers handled, the technique for each (removal, masking, consistent pseudonymous tokens, date shifting), the detector used on free text, and the residual-risk test on a reviewed sample. Under HIPAA, Safe Harbor removes 18 listed identifiers, while Expert Determination relies on a qualified expert's documented analysis; neither eliminates all re-identification risk [7]. The de-identification evidence package checklist lists the documents in detail.

A request template for the evidence pack

Sending one structured request avoids weeks of partial answers. The template below can go out with your evaluation license or NDA.

Illustrative example: invented to show structure; it does not describe an available dataset.

evidence_request:
  dataset_name: "B2B support tickets, English, 2021-2025"
  revision_requested: "tag or manifest SHA-256 of the files to be delivered"
  due_before: "contract signature"
  artifacts:
    sample:
      size: 2000 records
      draw_method: "random, seed stated, stratified by product_line and year"
      license: "evaluation only, delete on request"
    data_dictionary:
      per_field: [name, type, nullable, allowed_values, units, timezone, derived_from]
    profiling_report:
      computed_on: "full delivery population"
      metrics: [row_count, null_rate, cardinality, date_range, language_mix,
                exact_dup_rate, near_dup_rate_minhash_0.8]
    provenance_summary:
      fields: [source_system, collection_period, holder_legal_basis,
               consent_or_contract_reference, third_party_content_present]
    deidentification:
      fields: [identifiers_handled, technique_per_identifier, free_text_detector,
               residual_hit_rate_on_reviewed_sample, method_owner]
    labeling_qa:
      fields: [guideline_version, annotator_qualification, agreement_metric,
               adjudication_rate]
    test_results:
      tied_to: revision_requested
      correction_process: "how errors are reported, fixed and reissued"

The guide to requesting a training data sample covers sample sizing and handling, and evaluation licenses and NDAs for dataset samples covers the paper that governs it.

Red flags in the evidence you receive

Most evidence problems show up as mismatches between artifacts rather than missing ones. Treat these as reasons to pause signature:

  • The profiling report row count or date range does not match the sample's metadata or the quote.
  • Statistics were computed on the sample, or on an earlier revision, rather than the deliverable.
  • The data dictionary lists fields the sample lacks, or the sample has undocumented fields.
  • De-identification is described only as "anonymized" with no technique, detector or residual test.
  • Agreement metrics are reported without the number of doubly labeled items.
  • Provenance names a category ("enterprise customers") instead of a source system and legal basis.
  • Test results carry no revision identifier, so they cannot be tied to what ships.

Map each red flag to a contract response: a fix before signature, an acceptance criterion with a remedy, or a documented accepted risk. Acceptance criteria for licensed training data shows how to carry the thresholds into delivery.

How SourceX handles the evidence pack

SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records, documents, and finance and legal workflows, and manages the commercial process through licensing. Each dataset is rights-reviewed for ownership and consents, and diligence materials covering source, rights, preparation and allowed use are prepared per dataset. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. If you are assembling an approval file for a specific dataset, you can describe the data you need to SourceX.

Get the evidence for your data request

SourceX looks for US businesses that hold the data you describe, and every release is approved by the supplying company; nothing is contracted until a supplier agrees. Datasets are delivered under a license defining records, uses, term and delivery, through private, access-controlled workflows after an executed agreement. Start by describing the dataset and the evidence your approval file needs at SourceX for buyers.

Frequently asked questions

Should the profiling report come from the vendor or from us?

Ask the vendor for a report on the full population, then reproduce a subset of the metrics on the sample yourself. Agreement between the two is the evidence; either one alone is weaker.

How large should the evidence sample be?

Large enough to estimate the rates you care about. To check a duplicate or PII rate near 1%, a few hundred records gives a wide interval, while a few thousand gives a usable one.

What if the vendor will not release a sample before signature?

Offer an evaluation license with deletion terms, or ask for a clean-room review in the vendor's environment. If neither is acceptable, record the gap as an accepted risk and tie payment to acceptance criteria.

Sources

  1. SoftwareSeni, "How to require evaluation artifacts from AI vendors before signing any contract". https://www.softwareseni.com/how-to-require-evaluation-artifacts-from-ai-vendors-before-signing-any-contract
  2. Hugging Face Forums, "Is selling datasets way harder than building them? Or is it just me?". https://discuss.huggingface.co/t/is-selling-datasets-way-harder-than-building-them-or-is-it-just-me/177440
  3. Codebridge, "AI vendor evaluation checklist for accounting firm COOs". https://www.codebridge.tech/articles/ai-vendor-evaluation-checklist-for-accounting-firm-coos
  4. FISD Alternative Data Council, "Data Provider Due Diligence Questionnaire (DDQ) with GenAI Questions" (2024). https://fisd.net/wp-content/uploads/2024/02/FISD-Alternative-Data-Council-Due-Diligence-Questionnaire-with-GenAI-Questions-022824.docx
  5. Pushkarna, Zaldivar, Kjartansson (Google Research), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  6. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Data quality for analytics and ML, Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
  7. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  8. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data