Skip to content

Procurement, samples and ongoing supply

Questions to ask a training data vendor before you shortlist them

Quick answer

The questions to ask a training data vendor on a first call are the ones whose answers can disqualify it: who holds the rights and through which agreement, which systems or collection process produced the records, what notice or consent covered the people in them, whether a representative sample can be released and on what terms, how data and updates are delivered, and what unit the price is quoted in. Vague answers on rights or consent end the conversation. Specific ones earn a written questionnaire.

By SourceX Editorial · Updated

What a first call can settle, and what to leave for written diligence

A first call should establish whether the vendor can answer questions on rights, source, consent, samples, delivery and price specifically enough to justify a sample request; it is not the place to collect evidence. Contracts, consent records and test results come later, through a written due diligence questionnaire (DDQ), a sample evaluation and a dataset-level legal review.

First ask which kind of vendor you are talking to, because the follow-up questions depend on it:

  • Data holder: the company whose systems created the records. Rights questions go to its own customer and employee terms.
  • Broker or reseller: licenses data obtained from others. Ask for the upstream sublicensing language; see what to verify when buying through a broker or reseller.
  • Collection or annotation vendor: commissions new recordings or labels. Ask about participant agreements and IP assignment.
  • Synthetic data vendor: generates records with models. Ask which models and seed data.

The formal follow-up belongs in a data provider due diligence questionnaire and SourceX's AI training data due diligence checklist.

The first-call question set and the answers that disqualify

Fourteen questions fit a 30-minute call; the last column lists answers that should end or escalate the conversation. Ask them in order, because a failed rights answer makes the rest moot.

#AskA specific answer namesEscalate or walk away if
1Who holds the rights to these records, and what is your role?The entity that created or collected them; the vendor's role as holder, licensee, broker or collectorNo named holder, or "aggregation" without upstream contracts
2Which agreement lets you license this for model training, and does it allow sublicensing?The holder's customer or employee terms, an upstream license with sublicensing, contributor agreements, or a named open license"It's publicly available," or an upstream grant limited to analytics or internal use
3Was any portion scraped, crawled or taken from public datasets?Domains or datasets, crawl dates, robots.txt and site-terms handling, each dataset's original license"The open web," or license labels copied from a hosting page
4Was any portion generated or rewritten by a model?The generating model, its output terms, the synthetic share and the seed dataThe vendor does not track which records are synthetic
5Which systems produced the records, over what dates?Source applications (ticketing, CRM, EHR, call platform, code host), export method, date range, organization count, a data dictionary"Proprietary sources" with no field list
6What was filtered, de-duplicated, redacted, translated or labeled after export?Each step and tool, applied identically to the sampleA sample "cleaned up" for demonstration
7What notice or consent covered the people in these records, and does it cover third-party AI training?The notice or consent text and version date; recording-consent and biometric-release practicePermission added by a later terms change without new notice
8Is personal or health data present, and what de-identification method was applied?The method (for health data, HIPAA Safe Harbor or Expert Determination), the tool, a residual-identifier check"Anonymized," with no method named
9Can you release a sample, how is it selected, and on what terms?A random or stratified draw from the population you would buy, its size, NDA or evaluation-license termsOnly a hand-picked showcase
10Which quality claims can you state as numbers?Schema validity rate, exact and near-duplicate rates, audited label error rate, and the revision measured"High quality" or "good for fine-tuning"
11Who else has licensed this data, and has any of it been published?Number and type of other licensees, exclusivity options, any subset in public benchmarks or papersNo answer, for data meant for evaluation
12How will data, updates and deletions reach us?Format (JSONL, Parquet, WAV with a manifest), transfer route (cross-account bucket, Delta Sharing, warehouse share), versioning, refresh cadenceEmail attachments, or no way to propagate deletions
13What is the pricing unit, and what counts as one?Per record, per token (and tokenizer), per audio hour, per document or flat subscription; counted before or after de-duplicationA price with no unit
14What may we disclose about the source in our training-data documentation?Publishable fields: source type, owner, licensed status, personal information, collection period, synthetic shareConfidentiality that blocks any description of the source

Rights basis: the questions to settle before anything else

The rights questions come first because no sample or price can fix a vendor that cannot show a chain from the record holder to you. Alternative-data buyers ask at this depth: a Lowenstein Sandler guide says investment-firm buyers want detailed provenance covering paid subscriptions, surveys, industry conversations and web scraping, plus any past, current or threatened litigation or enforcement [1]. The FISD Alternative Data Council's illustrative DDQ, whose 2024 edition adds generative AI questions, asks whether a vendor buys data from others and, if so, for the contract terms that permit resale [2].

"Publicly available" is not a rights basis. The Data Provenance Initiative's audit of more than 1,800 text datasets reported license omission rates above 70% and error rates above 50% on popular dataset hosting sites [5], so an inherited license label needs tracing to its origin. Web permissions also shift: the Consent in Crisis authors report that between 2023 and 2024 about 5% of all C4 tokens, and over 28% of its most actively maintained critical sources, became fully restricted from use [6].

For model-generated records, ask about the generator's output terms; Anthropic's help center, for example, states that its terms do not allow outputs to be used to train competing models [7]. For open licenses with use limits, see how non-commercial datasets interact with commercial training.

Source systems and the documentation a vendor should already have

A vendor that knows its data can name the source application, export path and date range, and can send a data dictionary after the call. Ask what the records capture at the process level, not the industry level: a vendor guide to off-the-shelf training data itself notes that such data is generally oriented to foundational business areas rather than specific processes [4], so "customer service conversations" may lack the escalation, refund or tool-call steps your agent needs.

Ask what documentation exists. A datasheet in the sense of Gebru et al. records motivation, composition, collection process and recommended uses [8], and Croissant-RAI adds machine-readable responsible-AI fields to the Croissant metadata format [9]. A vendor with neither can still qualify, but you will write that documentation yourself.

Questions 7 and 8 decide whether privacy review can clear the data at all, so ask them before anyone loads a sample. The core test is whether the people in the records were told about this use at collection. FTC staff warned in 2024 that adopting more permissive data practices, such as AI training or third-party sharing, through a surreptitious, retroactive change to terms or privacy policies may be unfair or deceptive [10]. Remedies can reach models: the FTC's 2021 Everalbum order required deletion of models and algorithms developed from users' photos and videos [11].

Listen for answers tied to the governing law:

  • Call and meeting audio: California Penal Code section 632 prohibits recording a confidential communication without the consent of all parties [12]. See the call-recording consent checks for buyers.
  • Faces and voices: Illinois's Biometric Information Privacy Act covers voiceprints and face-geometry scans and requires written notice and a written release before collection [13].
  • Health records: HHS describes two HIPAA de-identification methods, Safe Harbor (removing 18 listed identifiers) and Expert Determination [14]. The vendor should name one.

The FISD template asks for copies of consent or other terms agreed with the individuals whose data is collected [2]. Request the notice text in writing and check it against the guide to consent and notice records.

Turning quality talk into numbers you can test

Accept only quality claims that a sample can confirm or refute. A Hugging Face forum discussion on selling datasets draws the line: "the JSONL is valid" or "the exact-duplicate rate is below 1%" can be checked, while "good for fine-tuning" cannot [3]. The thread also separates three claims: that a specific revision was tested, that the producer runs a repeatable correction process, and that the data was tested for a named buyer's intended use [3]. Ask which of the three the vendor can document.

Ask about three measures by name:

  • Duplicate rates. Lee et al. found many near-duplicates in common language-modeling datasets; models trained on deduplicated data emitted memorized text about ten times less often [15].
  • Audited label error rate. Northcutt et al. estimated an average label error rate of at least 3.3% across the test sets of ten widely used datasets [16], so ask how any vendor figure was measured.
  • Prior exposure. For evaluation data, ask whether any subset has been published or shared. In a February 2026 post, OpenAI said it no longer reports SWE-bench Verified because the benchmark is increasingly contaminated [17].

Then request a sample sized for those tests under evaluation license or NDA terms that permit them.

Delivery mechanics and the pricing unit

Format, transfer route and billable unit can change a cost comparison more than the headline number, so pin them down before any quote. JSON Lines files must be UTF-8 with one valid JSON value per line [18]; a vendor that says "JSON" may mean one multi-gigabyte array your loader cannot stream. For multi-terabyte deliveries, check that the plan uses a service still on offer: as of October 2026, AWS no longer offers Snowball Edge to new customers and points them to DataSync, Data Transfer Terminal or partner solutions [19].

Per-token prices depend on which tokenizer counts; per-record prices depend on whether duplicates and filtered records are billable. Get both answers, then normalize offers to cost per usable record and check them against common data license pricing structures.

Disclosure duties that make vendor answers part of your records

If you build or modify models, several laws can require you to describe training data publicly or to deployers, depending on the type of model and where it is offered, and your description will rest on what vendors tell you. As of October 2026:

  • California AB 2013 required developers of generative AI systems made available to Californians to post training-data documentation by January 1, 2026, covering items such as dataset sources or owners, copyrighted or licensed material, personal information, collection periods and synthetic data [20]. See the records to collect for AB 2013.
  • EU AI Act Article 53(1)(d) requires general-purpose AI model providers to publish a sufficiently detailed training-content summary using the AI Office template, an obligation in effect since August 2, 2025 [21].
  • Colorado SB26-189 requires developers of automated decision-making technology that materially influences consequential decisions to give deployers documentation including training data categories, with its main duties starting January 1, 2027 [22].

Question 14 tests whether a vendor's confidentiality terms let you meet these duties.

When an intermediary sources the data for you

With an intermediary, ask the same questions about the underlying record holder, plus two more: did the holder approve this release, and under what agreement does the intermediary act? SourceX is one such intermediary: it sources operational datasets from US companies on request, buyers describe the data rather than the businesses, and every release is approved by the supplying company. Each dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place, and diligence materials on source, rights, preparation and allowed use are prepared per dataset for the buyer's review. A request does not guarantee a matching dataset; you can describe the data you need on SourceX's buyer page.

After the call: what to put in writing

Send a written recap within a day and ask the vendor to correct it; a spoken "training is fine" is not a license term. Move vendors that pass to the written DDQ, collect the evidence to request from a data vendor before you sign, and rank them on the data vendor evaluation scorecard. SourceX's guide to evaluating data supplier quality and the AI training data procurement hub cover the remaining steps.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Qualifying a data supplier for your next purchase?

At SourceX's buyer page you can submit a data request, talk to SourceX or browse the dataset catalog. Describe the records you need; SourceX looks for US businesses that hold them, checks the data and the supplier's licensing permissions, and manages the license and delivery, and nothing is contracted until a supplier agrees. See how SourceX works with data buyers.

Sources

  1. Lowenstein Sandler LLP, "Key considerations for alternative data and AI vendors to investment firms". https://www.lowenstein.com/media/iyrpwxij/key-considerations-for-alternative-data-and-ai-vendors-to-investment-firms.pdf
  2. FISD Alternative Data Council, "Data Provider Due Diligence Questionnaire (DDQ) with Generative AI Questions" (2024). https://fisd.net/wp-content/uploads/2024/02/FISD-Alternative-Data-Council-Due-Diligence-Questionnaire-with-GenAI-Questions-022824.docx
  3. Hugging Face Forums, "Is selling datasets way harder than building them? Or is it just me?". https://discuss.huggingface.co/t/is-selling-datasets-way-harder-than-building-them-or-is-it-just-me/177440
  4. Shaip (vendor page), "Choosing the right off-the-shelf AI training data provider". https://www.shaip.com/blog/choosing-the-right-off-the-shelf-ai-training-data-provider
  5. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version in Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
  6. Longpre et al., "Consent in Crisis: The Rapid Decline of the AI Data Commons" (2024). https://arxiv.org/pdf/2407.14933
  7. Anthropic, "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
  8. Gebru et al., "Datasheets for Datasets" (2021). https://arxiv.org/pdf/1803.09010
  9. Jain et al., "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  10. Federal Trade Commission, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  11. Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
  12. California Legislature, "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632
  13. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  14. U.S. Department of Health and Human Services, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  15. Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  16. Northcutt et al., "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  17. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  18. jsonlines.org, "JSON Lines". https://jsonlines.org/
  19. Amazon Web Services, "AWS Snowball Edge availability change" (2025). https://docs.amazonaws.cn/en_us/snowball/latest/developer-guide/snowball-edge-availability-change.html
  20. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  21. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  22. Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data