Skip to content

Evaluation and benchmarking datasets

How to commission a custom LLM evaluation dataset

Quick answer

To commission a custom LLM evaluation dataset, fix the capability, failure modes and scoring rubric before you contact providers, then decide where the items come from: experts writing new ones, licensed business records whose recorded outcomes supply the labels, or both. Put the deliverables and the held-out controls into a statement of work, pay for a pilot batch before production, and accept the set only after your own blind audit of its gold labels.

By SourceX Editorial · Updated

General purchasing steps are in how to procure enterprise training data; the LLM evaluation datasets hub maps the alternatives to commissioning.

What a commissioned eval set should contain when it arrives

A commissioned evaluation set is a versioned package of held-out inputs, a gold label per item with a record of how it was produced, slice tags, adversarial items and documentation. Often your own team, not the provider, runs the evaluation.

Vendor material shows the same shape. AIXBlock describes building held-out inputs, gold labels, slices and red-team prompts while the buyer's ML team runs the evaluation [1]; AfterQuery scopes custom evaluation datasets around a capability, workflow, domain, rubric or evaluation target [2]; Dataforce argues that a test set never published online avoids the risk that models have already trained on it [3]. These are self-descriptions of market practice, not evidence of quality.

The unit you buy is the item record. Ask for one JSON Lines record per item plus a manifest of item IDs and SHA-256 hashes per set version, so you can prove which version a model saw.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "item_id": "rf-0417",
  "set_version": "2026-09-r1",
  "split": "test",
  "task": "refund_policy_decision",
  "input": {
    "messages": [{"role": "user", "content": "Two of my three mugs arrived cracked and I used a 20% coupon. What do I get back?"}],
    "context_docs": ["policy/refunds-v7.md#partial-returns"]
  },
  "gold": {
    "answer": "Refund two mugs at the coupon-adjusted unit price; no return shipping charge.",
    "acceptable_variants": ["Pro-rated refund for the two damaged units, free return label"],
    "evidence_spans": ["policy/refunds-v7.md#L41-L48"],
    "rubric_id": "refund-decision-v3"
  },
  "label_provenance": {
    "method": "expert_written_independent_review",
    "author_id": "w-12",
    "reviewer_ids": ["r-07"],
    "adjudicated": false,
    "rationale": "Coupon discount is applied pro rata per policy section 4.2."
  },
  "slices": ["partial_return", "coupon_applied"],
  "adversarial": false,
  "source": {"origin": "expert_written", "created_at": "2026-09-18"},
  "license_scope": "evaluation_only"
}

Two fields carry most weight at acceptance: label_provenance, showing how each gold label was produced, and created_at, which isolates items written after a model's training cutoff (post-cutoff evaluation data).

Four decisions to make before you contact a provider

You cannot compare proposals until you have fixed what the set must decide, which failures it must catch, how items are scored and how many items each slice needs. Otherwise providers fill the gaps with whatever is fastest to produce.

  1. The decision. Choosing between two models needs a paired comparison on identical items; gating a release needs pass criteria per slice (when a public benchmark is enough).
  2. Failure taxonomy and slices. Turn each failure you have seen or fear (wrong policy branch, unsupported citation, missed escalation, malformed tool call) into a slice with a minimum item count, over-allocating rare, high-risk cases (stratified evaluation sets).
  3. Scoring method. Exact match or programmatic checks where an objective answer exists, human rubric grading where it does not, an LLM judge only after calibration against human labels. The LiveBench authors chose objective ground-truth scoring partly because, they say, LLM judges can introduce biases and break down on hard questions [4].
  4. Items per slice. Work back from the smallest difference you must detect. A 2025 ICML position paper argues that CLT-based confidence intervals are too narrow for LLM evaluations with fewer than a few hundred items [5]; see sizing an eval set for statistical power and what a custom evaluation dataset costs.

Put the answers in a one-page scope memo with a draft rubric and a few items you labeled yourself. It seeds the full evaluation dataset specification and becomes the yardstick for the pilot.

Where the items come from: expert-written, licensed records or a hybrid

Choose the item source by what your ground truth must be. Expert-written items suit capabilities judged against a rubric; licensed business records suit tasks where a recorded decision or outcome can serve as the label; a hybrid adds expert labels to real inputs that lack one.

Expert-written itemsLicensed business recordsHybrid: real inputs, expert labelsGenerated items, human-checked
Source of the gold labelWriter plus independent reviewerRecorded outcome: final ticket disposition, merged fix, adjudicated claimExperts labeling real inputs against your rubricA model, verified by people
Realism of inputsDrifts toward clean, templated promptsReal formats, noise and edge casesRealLow to medium
Contamination exposureLow if never published or reusedLow if the records were never publicLowCan echo public items the generator saw
Main rights questionWho owns the itemsWhether the record holder may license them, and consent limitsBothGenerator terms of use
Typical failureItems easier than production trafficOutcomes encode the business's own errorsExpert and outcome disagree, unresolvedItems mirror one model family's style

SWE-Bench Pro shows a version of the business-records route in coding: next to public and held-out sets from copyleft repositories, its authors added a commercial set built from startup codebases under partnership agreements, whose results are published while the code stays private [6]. The same logic applies to support tickets, claims and contracts (outcome-labeled evaluation data, golden datasets from business records); generated items carry their own risks (synthetic evaluation data limits).

When inputs should come from real business work, SourceX sources operational datasets from US companies, such as support and sales histories, engineering records, and finance and legal workflows, and manages the licensing agreement. Every release is approved by the supplying company, and datasets are sourced on request rather than held in stock, so a request does not guarantee a match. See evaluation datasets built from real business work, or tell SourceX which records and outcomes you need.

Statement of work clauses specific to evaluation data

An eval-set statement of work (SOW) differs from a training-data SOW in three places: it specifies how gold labels are produced and checked, binds the provider to held-out controls, and defines acceptance by label correctness rather than volume. Start from the general statement of work for custom data collection and add these clauses.

SOW clauseWhat to specifyWhy it matters for evaluation
Item specificationTask types, input formats, context documents, slice quotas, share of adversarial itemsStops drift toward whatever is fastest to write
Writer and rater qualificationsCredentials or role experience per slice; conflict declarationsGold labels are only as good as their authors (sourcing domain-expert raters)
Gold-label methodIndependent second label on a stated share of items, an adjudication rule, an agreement statistic reported per sliceKrippendorff's alpha measures agreement among raters labeling the same items; 1 is perfect reliability, 0 its absence [7]
Rationale and evidenceA short rationale and evidence span for every gold labelAuditors check labels without re-deriving them
Rubric controlRubric version on every item; a rubric change triggers re-labeling of affected itemsMixed rubric versions make scores incomparable (rubric design with domain experts)
Measurement definitionsEach acceptance measure written as a property plus its measurement methodMirrors how ISO/IEC 5259-2 defines a quality measure element [8]
Drafting toolsWhether writers may use LLMs, which ones, under what retention termsItems pasted into hosted models leave your perimeter
DocumentationA datasheet covering motivation, composition, collection process and recommended uses [9], also as Croissant-RAI JSON-LD [10]Reviewers who never see items can still judge the set
Defect remedyReplacement of items that fail audit at the provider's costDefines what "accepted" means

Machine-readable documentation is becoming an expectation: NeurIPS 2026 sets Croissant-RAI-based responsible-AI metadata requirements for its Evaluations and Datasets Track [11].

Held-out controls that keep commissioned items out of training data

A commissioned set stays useful only while no model you test has trained on its items, paraphrases or near-duplicates. Leaks come through the provider's other business, its writers' tools and your own evaluation runs, so contract for controls.

Public benchmarks show the stakes. In a post dated 23 February 2026, OpenAI said it had stopped reporting SWE-bench Verified because it was increasingly contaminated, so score gains increasingly reflected training-time exposure [12]. A 2024 study of benchmark inflation notes that a private holdout could verify published scores, but most benchmarks lack one [13]. A commissioned set is that holdout; its secrecy is what you pay for.

Controls to write into the SOW and license:

  • Provider conflicts. The provider discloses whether it sells training data, annotation or evaluation services to developers of models you will test. A 2025 paper on private evaluation curators describes this conflict of interest and the transparency lost behind closed doors [14]; see third-party evaluation vendor independence.
  • No reuse, including paraphrases. Items, paraphrases and translations may not enter any dataset the provider sells. LMSYS researchers showed a 13B model trained on rephrased MMLU test items scored 85.9 on MMLU while n-gram overlap checks failed to flag it [15].
  • Segregated storage and a named access list, with access logs available on request.
  • Writer tooling. No pasting items into consumer chatbots or hosted models without approved no-retention terms.
  • Canary marker. A unique string embedded in every item file so copies can be searched for later.
  • Deletion certificate for provider working copies after acceptance.
  • Contamination screen at delivery against the benchmarks and training sets you use, with embedding or paraphrase checks, not only verbatim n-gram matching (contamination-resistant evaluation design).
  • Your own exposure log of which model and API saw which version (keeping a private eval set private).

Put the same questions to every provider. For reference, SourceX does not train AI models and delivers through private, access-controlled workflows, never email attachments.

A staged commissioning process with a paid pilot

Commission in stages so production starts only after a pilot batch shows that the provider's writers and raters reproduce your labels.

StepOwnerOutputGate to the next step
1. Scope memoEvaluation leadDecision, taxonomy, scoring, slice sizes, draft rubricProduct owner sign-off
2. Request and shortlistProcurement and evaluation leadRequest or RFP with the eval-specific clauses (RFP template, writing a data request)Conflict disclosures received
3. NDA and pilot termsCounsel, privacy, securityConfidentiality, pilot scope, pilot item ownership (evaluation licenses and NDAs)Signed
4. Pilot batchProviderItems in every slice, with rationales and provenanceAgreed format
5. Blind re-labelYour domain expertsIndependent labels on a random pilot sampleAgreement at or above the SOW threshold in every slice
6. Rubric revisionEvaluation lead with providerRevised rubric; re-labeled pilot itemsDisagreements explained, not just counted
7. Production in tranchesProviderTranches with hashed manifestsEach tranche passes a sampled audit
8. Acceptance auditYour auditorsDefect list and replacementsDefect rate within the SOW limit
9. Freeze and handoverSecurityFrozen version in restricted storage; deletion certificateExposure log opened
10. RefreshEvaluation leadReplacement items for saturated or leaked slicesOn a set schedule (refresh cadence)

If items derive from records with personal data, de-identification must keep what makes each item hard (de-identifying evaluation data without breaking the test).

Accepting the set: auditing gold labels before sign-off

Accept a commissioned set on gold-label correctness measured by your own blind audit, not on item counts or the provider's reported agreement. Even well-known public test sets carry wrong answers: Northcutt and colleagues estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets and showed that such errors can change which model ranks higher [16]. A private set has no outside community to find its errors for you.

Define defects in the SOW so acceptance is not a negotiation: a wrong gold label, an item ambiguous under the rubric, a wrong slice tag, a near-duplicate, an item found in a public source, or personal data left in place. Audit a random sample per slice, adjudicate disagreements with a third expert, and have failed items replaced; the sampling plan is in accepting a delivered eval set.

License and ownership terms that differ for evaluation data

An eval set's license must protect secrecy and comparability, not just define permitted use. Settle these before production:

  • Ownership or exclusivity: whether you own the items or hold an exclusive license, and whether the provider may sell items written to the same rubric.
  • Use scope: evaluation only, or also training once the set is retired; say so up front if you may want the second.
  • Hosted model APIs: whether items may be sent to third-party models during evaluation, and under what retention terms.
  • Publication: whether you may publish scores, sample items or the rubric (publishing results on licensed evaluation data).
  • Underlying rights: for items built from business records, the holder's permission to license them and any consent limits (evaluation-only data license terms).

For records sourced through SourceX, rights review checks that the business owns or may share them and that required consents are in place.

Mistakes that make a commissioned eval set unusable

  • Letting the provider write the rubric, so the set measures its idea of quality.
  • Gold labels written by the item's author with no independent review.
  • Accepting an overall agreement figure while one high-risk slice is poor.
  • Using slices of a few dozen items to decide small differences between models, with CLT error bars that understate the uncertainty [5].
  • Undisclosed LLM drafting, so you cannot rule out that items mirror one model family's phrasing.

Need evaluation items drawn from real business records?

Describe the task, the records or outcomes that should supply gold labels, the slices and the license scope you need. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Describe your evaluation data needs.

Sources

  1. AIXBlock (vendor blog, market practice), "LLM Evaluation Datasets: Held-Out Sets for Production". https://www.aixblock.io/blogs/llm-evaluation-datasets
  2. AfterQuery (vendor page, market practice), "Buy AI training data". https://www.afterquery.com/buy-ai-training-data
  3. Dataforce (vendor blog, market practice), "Disadvantages of standard LLM benchmarks". https://www.dataforce.ai/blog/disadvantages-standard-llm-benchmarks
  4. White et al. (arXiv:2406.19314; ICLR 2025), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
  5. ICML 2025 (poster), "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  6. arXiv:2509.16941, "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/html/2509.16941v1
  7. Klaus Krippendorff, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  8. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
  9. Gebru et al. (arXiv:1803.09010; Communications of the ACM 2021), "Datasheets for Datasets" (2018). https://arxiv.org/pdf/1803.09010
  10. Jain et al., MLCommons Croissant RAI task force (arXiv:2407.16883), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  11. NeurIPS blog, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
  12. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  13. arXiv:2410.09247, "Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts" (2024). https://arxiv.org/pdf/2410.09247.pdf
  14. arXiv:2503.04756, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
  15. LMSYS Org, "Catch me if you can! How to beat GPT-4 with a 13B model" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
  16. Northcutt, Athalye, Mueller (arXiv:2103.14749; NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data