Skip to content

Evaluation and benchmarking datasets

Third-party eval data vendors: independence, conflicts and verification

Quick answer

The main risk with a private evaluation data curator is that you cannot see the test set, so you are trusting the vendor's independence instead of checking it. That trust breaks down when the same vendor also sells training data in the same domain, reuses items across clients, or runs the scoring itself. Vet an eval vendor on four things: commercial conflicts, item provenance and separation, label methodology with agreement statistics, and whether you can reproduce the scores on your own infrastructure.

By SourceX Editorial · Updated

Why private eval curators carry a structural conflict

A vendor that sells training data to model developers and also builds or runs the hidden test set those developers are judged on has a conflict of interest, even when nobody acts in bad faith. Bansal and Maini make this argument about private data curators, noting that open benchmarks give full transparency while private ones ask users to trust the curator's process [1]. Their paper also flags a subtler issue: the same annotator pool, guidelines and stylistic preferences can shape both a vendor's training data and its grading, so models trained on that vendor's data may score well partly because they share its taste [1].

This combination exists in the market. Some data vendors advertise custom training data and custom evaluation datasets from one catalog [2]. That is not disqualifying, but it moves the burden of proof onto the vendor. Your procurement question is not "is this vendor honest" but "what evidence shows the eval items never touched any training pipeline, ours or another client's?"

The conflict matters most in three situations:

  • Model selection. You are comparing foundation models, and one of them was fine-tuned on data from the same vendor that wrote your test set (see running a model bake-off with your own eval set).
  • Your own training data. You buy training data and eval data in the same domain from the same supplier, so leakage would inflate your own scores.
  • Vendor-run leaderboards. A vendor publishes rankings on a hidden set while selling data to some of the ranked labs.

How contamination turns a vendor score into a training artifact

When eval items or close paraphrases leak into training data, score gains stop measuring capability and start measuring exposure. OpenAI stopped reporting SWE-bench Verified in February 2026, stating that contamination meant improvements increasingly reflected training-time exposure to the benchmark [4]. LiveBench's authors describe the same failure for static test sets and respond by refreshing questions over time [8].

A private vendor set is not immune; it simply moves the leakage path from the public web to the vendor's own operations. Typical leak routes include:

  • An item written for an eval engagement gets recycled into a training "seed set" for another client.
  • Annotators who wrote gold answers later produce SFT or preference data in the same domain from memory, with templates and phrasings intact.
  • The vendor sends eval prompts to a hosted model API for scoring, and the provider's data retention policy allows those prompts to be used later.
  • Near-duplicates survive because the vendor deduplicates only on exact hashes, not on n-gram or embedding similarity.

Controls for the first and last routes belong in your contamination-resistant evaluation design, and the access-control side is covered in keeping a private eval set private. Contamination testing on delivered sets is covered in the owner guide on contamination checks for licensed evaluation data.

Separation controls to require in writing

Ask for separation controls the vendor can attest to and you can audit, not general assurances. The strongest set combines organizational, technical and contractual controls.

Organizational. Different teams, and ideally different annotator pools, for eval items and training data in the same domain. Ask how the vendor prevents an annotator from moving between an eval project and a training project for the same capability within a set cooling-off window.

Technical. Every eval item should carry a unique ID, a creation timestamp and a content hash, stored in a registry the training-data side cannot write to. Ask whether the vendor screens outbound training deliveries against that registry using exact hashes, MinHash or n-gram overlap, and embedding similarity, and what thresholds it uses. Canary strings embedded in eval files let you detect later ingestion.

Contractual. An item non-reuse attestation: the vendor warrants that items delivered to you were not previously sold, will not be resold, and will not be used to create training data for anyone. Pair this with an evaluation-only license on the buyer side so obligations run both ways, plus an audit right over the item registry.

Verifying scores: run grading where you control it

The cleanest way to reduce dependence on a vendor is to take delivery of the items and gold labels and run scoring yourself. Some vendors are built around this model, delivering held-out sets that the buyer evaluates on its own infrastructure [3]. If the vendor must host the set, for example to keep it out of your fine-tuning pipelines, ask for per-item outputs, grader prompts and rubric versions so you can re-score a sample independently.

Treat a vendor-reported score as a claim to reproduce. A minimal verification run looks like this:

  1. Pull a stratified sample of 100 to 300 items across slices (see sizing an eval set for statistical power for how large it must be to detect the deltas you care about).
  2. Re-run the model with the exact decoding settings the vendor used (temperature, max tokens, system prompt, tool configuration).
  3. Re-grade with your own grader, either human reviewers or a pinned LLM judge with a versioned rubric.
  4. Compare per-item pass/fail, not just aggregate accuracy. Disagreement concentrated in one slice usually points to a rubric ambiguity or a mislabeled gold answer.

Label methodology and agreement statistics to demand

A vendor's gold labels are only as good as its labeling process, and published test sets show how far that can drift. Northcutt and colleagues estimated an average label error rate of at least 3.3% across test sets of 10 widely used datasets, including at least 6% of the ImageNet validation set [5]. On a small private set, a few percent of wrong gold answers can flip a model comparison.

Ask for the methodology as a document, not a sales summary. It should state rater qualifications (credentials, domain tenure, screening tests), the number of independent labels per item, adjudication rules, and inter-rater agreement per slice. Krippendorff's alpha suits this because it handles multiple raters, missing ratings and nominal, ordinal or interval scales [6]; Cohen's kappa is acceptable for two-rater designs. Then run your own gold-label audit on the delivered set before you trust it.

Red flags include a single annotator per item with no adjudication, agreement reported only as raw percent, agreement computed on an easy calibration batch rather than production items, and rubrics that changed mid-project without versioning.

Eval vendor due-diligence questionnaire

Adapt a general data-provider DDQ, such as the FISD Alternative Data Council template, which added generative AI questions in its 2024 edition [7], with eval-specific sections. SourceX's data provider due diligence questionnaire covers rights, security and delivery; the questions below cover independence.

Illustrative example: invented to show structure; it does not describe an available dataset.

SectionQuestionEvidence to requestRed flag
Commercial conflictsDo you sell training, SFT or preference data in this domain? Do you sell to any model developers we will rank?Written disclosure by domain and client typeRefuses to disclose or answers only "we keep data separate"
Commercial conflictsDo you publish leaderboards on sets drawn from the same item pool?Leaderboard methodology and item-pool lineageSame pool used for public rankings and private client evals
Item provenanceHow was each item created: written from scratch, adapted from public sources, or drawn from client records?Per-item source_type field and source referencesItems paraphrased from public benchmarks without disclosure
SeparationWhich teams and annotators worked on this set, and on which training projects in the same domain?Org chart excerpt, annotator pool IDs, cooling-off policyShared annotator pool with no restrictions
SeparationHow do you screen training deliveries against eval items?Dedup method, similarity thresholds, last screening logExact-hash matching only
ReuseHave any of these items been delivered to another client?Signed item non-reuse attestationAttestation limited to "verbatim" reuse
LabelingLabels per item, adjudication rule, rater qualifications?Labeling guideline version, rater screening criteriaOne label per item, no adjudication
AgreementInter-rater agreement by slice?Krippendorff's alpha or kappa per slice with item countsRaw percent agreement only
ReproducibilityCan we run scoring ourselves?Item and gold file delivery in JSONL or Parquet, grader prompts, rubric versionsScores available only through vendor dashboard
Model exposureDo eval prompts ever pass through third-party model APIs?List of APIs, data retention and training-use settingsUnknown or default retention settings
RefreshHow are items retired and replaced as they leak or saturate?Retirement log and refresh cadenceNo retirement process

Score each section pass, conditional or fail, and route conditional answers to your supplier risk tiering process. Broader supplier quality questions sit in the owner guide on evaluating data supplier quality.

When to switch to independent or self-built eval data

If a vendor cannot answer the commercial-conflict and separation rows credibly, the practical fix is to change the source of the test set rather than negotiate harder. Options include commissioning a set from a vendor that does not sell training data in your domain, building a golden dataset from your own business records, or licensing real operational records from a company outside the model-training supply chain under evaluation-only terms. If you want operational records sourced for you, describe the evaluation data to SourceX in terms of records and fields rather than named companies; allowed uses are agreed in each license.

Sets built from real business outcomes, such as resolved support tickets or completed contract redlines, have a structural advantage: the gold label is what actually happened, not an annotator's preference, which reduces the shared-annotator bias the private-curator paper describes [1]. The tradeoff is de-identification work and rights review, both of which a reputable source should handle before delivery. Independent evaluation organizations face their own version of these questions, covered in data rights for independent AI evaluators. For the wider map of benchmarks and private sets, start at the evaluation datasets hub.

Sourcing independent evaluation data from operational records

SourceX sources operational datasets, such as support and sales histories, engineering records, and finance and legal workflows, from US companies on request and manages licensing and ongoing purchases. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, with personal details removed or replaced before delivery, and SourceX does not train models. Describe the evaluation data you need at SourceX for buyers.

Sources

  1. arXiv (Bansal and Maini; ICLR 2025), "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
  2. AfterQuery, "Buy AI training data". https://www.afterquery.com/buy-ai-training-data
  3. AIxBlock, "LLM Evaluation Datasets: Held-Out Sets for Production". https://www.aixblock.io/blogs/llm-evaluation-datasets
  4. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  5. arXiv (Northcutt, Athalye, Mueller; NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  6. University of Pennsylvania, Annenberg School for Communication, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  7. FISD Alternative Data Council, "Data Provider Due Diligence Questionnaire (DDQ) with GenAI questions" (2024). https://fisd.net/wp-content/uploads/2024/02/FISD-Alternative-Data-Council-Due-Diligence-Questionnaire-with-GenAI-Questions-022824.docx
  8. arXiv (ICLR 2025), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data