Skip to content

Evaluation and benchmarking datasets

What a custom evaluation dataset costs: drivers and budgeting

Quick answer

A custom LLM evaluation dataset costs whatever its five drivers multiply out to: the number of items your statistical target requires, the minutes a qualified expert needs to write and grade each item, how many raters see each item, how much real context (documents, tickets, tool state) each item carries, and how often the set must be refreshed. Published price tiers are estimates [4]. Build the budget bottom-up from these drivers, then compare quotes against it. For the wider landscape, see the LLM evaluation datasets buyer map.

By SourceX Editorial · Updated

Item count comes from a power calculation, not a round number

Item count is the first cost lever to fix, and it should be derived from the smallest score difference you need to detect. Required items rise steeply as the target effect shrinks [2]: for a pass/fail metric, the standard error is roughly sqrt(p(1-p)/n), so halving the width of a confidence interval needs about four times as many items. A set meant to separate two candidate models by a few points needs far more items than one meant to catch a regression from 90% to 70%.

Small sets also mislead in a second way. A 2025 ICML position paper argues that central-limit-theorem confidence intervals are unreliable below a few hundred datapoints and tend to be too narrow [3]. If your budget only covers 150 items, plan for bootstrap or Bayesian intervals and say so in the eval report, rather than reporting tight error bars you cannot defend.

Stratification multiplies count. If you need a reliable score on each of six slices (for example, refund policy, escalation, billing disputes, account security, multilingual tickets, and adversarial prompts), each slice needs its own adequate sample. See stratified evaluation sets for rare and high-risk cases for allocation methods that keep rare slices from dominating the budget.

Expert time per item usually outweighs everything else

Per-item cost is mostly the loaded hourly rate of the person who writes the reference answer, multiplied by their minutes per item. A generalist grading a short classification label might spend under a minute; a licensed tax preparer writing a reference answer to a multi-document question with citations might spend 30 to 60 minutes. Curating even a modest set of high-quality examples is labor-intensive, as the LIMA authors noted for their 1,000 hand-curated pairs [7].

Three task properties push minutes up:

  • Reference-answer authoring versus judgment. Writing a gold answer from scratch costs more than grading a model output against a rubric.
  • Source reading. Items grounded in long contracts, claim files or codebases require the expert to read the context first; see long-context evaluation on real documents.
  • Rubric depth. Multi-criterion rubrics with partial credit take longer to apply than binary checks. Invest once in rubric design with domain experts so per-item grading stays fast.

Sourcing scarce specialists (clinicians, securities lawyers, SAP functional consultants) also adds recruiting and qualification overhead. Sourcing domain-expert raters covers screening tests and calibration rounds, which belong in the budget as fixed costs.

Rater redundancy and adjudication set the quality ceiling

Every additional independent rater per item adds roughly one more unit of grading cost, plus adjudication time when raters disagree. Single-rater gold labels are cheapest but give you no inter-annotator agreement estimate (Cohen's or Fleiss' kappa, Krippendorff's alpha), so you cannot tell label noise from model error. A common pattern is double-rating a calibration subset, measuring agreement, and single-rating the remainder only if agreement clears a threshold you set in advance.

Budget an adjudication pass for disagreements and a gold-label audit at acceptance. Accepting a delivered eval set describes sampling plans for that audit. Curator independence matters too: research on private data curators flags conflicts of interest and annotator bias when the party that builds the eval also supplies training data to the developers being evaluated [5].

Context capture and source data change the cost structure

Items built from real operational records cost differently from items written from scratch. Commissioned items pay mainly for authoring time; items derived from licensed business records (support tickets with resolutions, closed insurance claims, merged pull requests) pay for licensing, de-identification and selection, then for expert labeling on top. The trade is realism and outcome ground truth versus control over coverage.

Deliverable components each carry cost: inputs, gold labels, slice tags, metadata and red-team prompts [1]. For agent evaluations, add environment state (database snapshots, tool mocks, file trees) and state-based graders; agent evaluation task suites lists the components. For RAG evaluation, each item needs a question, an answer and the supporting passage IDs; see question-answer-citation triples.

When the source is real records, extra line items appear: rights review, removal or replacement of personal identifiers, and, for health data, HIPAA de-identification by Safe Harbor or Expert Determination. If the AI system is high-risk under the EU AI Act, its testing data must also meet Article 10's quality and governance criteria [8], which adds documentation work; as of October 2026, Regulation (EU) 2026/1744 reportedly moved those obligations to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems. For how outcome labels from real decisions work as ground truth, see outcome-labeled evaluation data.

Refresh and contamination control are recurring costs

An evaluation set is a depreciating asset: once items leak into training data or teams tune against them, scores inflate. Budget for a refresh cadence (quarterly or per model release are common choices) and for a held-out partition that is never shown to model developers. Contamination-resistant evaluation design covers canary strings, temporal holdouts and access controls.

Refresh cost is usually lower per item than the first build, because the rubric, rater pool, tooling and spec already exist. It is not zero: new items still need expert time, and drift in product policy can force relabeling of existing gold answers. Time-stamped records created after a model's training cutoff are one way to refresh; see post-cutoff evaluation data.

Budget worksheet: build the estimate bottom-up

A bottom-up worksheet makes quotes comparable, because every vendor or internal team prices the same line items. Fill in your own rates; the example below uses invented quantities and leaves the hourly rate as a variable so it does not imply market prices.

Illustrative example: invented to show structure; it does not describe an available dataset.

Line itemDriverExample quantityFormula
Spec and rubric designSenior expert hours40 h40 x R_senior
Rater recruiting and calibrationRaters x calibration hours6 raters x 4 h24 x R_expert
Item authoringItems x minutes per item600 x 20 min200 x R_expert
Second rating (calibration subset)25% of items x minutes150 x 8 min20 x R_expert
AdjudicationDisagreement rate x items x minutes15% x 150 x 10 min~4 x R_senior
Source dataLicensed records, de-identification600 source recordsLicense fee + prep (quoted)
Acceptance auditSampled items x minutes60 x 10 min10 x R_senior
Refresh (per cycle)New items + relabel150 x 20 min50 x R_expert per cycle
Tooling and project managementPercentage of labor10-20%Labor subtotal x overhead

Use the worksheet in three passes. First, set item count from your power target and slice plan. Second, run a 20-item pilot to measure real minutes per item and disagreement rates, because guessed minutes are the most common budget error. Third, multiply by quoted rates and add a contingency for relabeling after the first audit.

Build in-house, commission, or license: how the cost profile shifts

The cheapest route depends on volume, domain breadth and whether you already have the experts. Annotation vendors themselves note that in-house work tends to win at low, steady volume in a narrow domain, while outsourcing suits spikier or multi-domain needs [6]. Treat vendor framing as directional.

RouteMain costHidden costFits when
In-house with your own expertsExpert salary timeDiverted product work; weaker independenceNarrow domain, steady cadence
Commissioned from a labeling vendorPer-item or per-hour quotesRater quality variance; contamination risk if the vendor also trains models [5]Broad coverage, fast ramp
Licensed real records plus expert labelsLicense fee, preparation, labelingRights review, de-identification, supplier approvalRealism and outcome labels matter
Synthetic generation plus human reviewReview timeDistribution mismatch; see synthetic evaluation data limitsEarly prototyping

Licensing terms move price as well. Scope (evaluation-only versus training), exclusivity, term, publication rights for benchmark results and refresh obligations are all negotiable; evaluation-only license terms and publishing results on licensed eval data cover them. For general price drivers of licensed enterprise data, see what drives the price of licensed enterprise data.

Questions to put in every quote request

A good quote request removes ambiguity about the drivers above, so the vendor's number reflects your actual spec. Ask each supplier to state:

  • Item count per slice and the statistical target it was sized for.
  • Rater qualifications, raters per item, and the agreement metric they will report.
  • Measured minutes per item from a pilot, not an estimate.
  • What each item contains (input, gold answer, rationale, citations, slice tags, environment state) and the file format (JSONL with a published schema is a reasonable default).
  • How gold labels will be audited and who pays for relabeling.
  • Refresh pricing per cycle and how held-out partitions are protected.
  • For real-record sources: who owns the data, how personal details were removed, and what uses the license allows.

Pair this list with a written evaluation dataset specification and the AI training data RFP template. The full commissioning workflow is in how to commission a custom LLM evaluation dataset.

When the eval set should be grounded in real operational records rather than written from scratch, SourceX sources operational datasets from US companies on request and manages licensing; you can describe the evaluation data you need.

Pricing evaluation data grounded in real business records

SourceX sources operational datasets, such as support histories, engineering records and finance or legal workflows, from US companies on request; nothing is held in stock, and a request does not guarantee a match. Each dataset is rights-reviewed, personal details are removed or replaced before delivery, and pricing and allowed uses are agreed per deal in a license. Describe the evaluation data you need.

Frequently asked questions

Why do quotes for the same eval set vary so widely?

Vendors usually assume different item counts, rater qualifications and redundancy levels. Normalize every quote to the worksheet line items, and ask for measured minutes per item from a pilot.

Can I save money by using an LLM as the grader?

An LLM judge cuts per-item grading cost but still needs a human-labeled calibration set to measure its agreement with experts, and judges can favor outputs from their own model family. Budget for that calibration set and for periodic re-checks.

Is a smaller expert-graded set better than a larger crowd-graded one?

Often, for specialist domains, because label noise from unqualified raters can exceed the difference you are trying to measure. Check that the smaller set still meets your power target, and use intervals suited to small samples [3].

Sources

  1. AIxBlock, "LLM Evaluation Datasets: Held-Out Sets for Production". https://www.aixblock.io/blogs/llm-evaluation-datasets
  2. Tian Pan, "Statistical power for LLM evals" (2026). https://tianpan.co/blog/2026/04/15/statistical-power-llm-evals
  3. ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  4. AI Superior, "Cost of private LLM evaluation services". https://aisuperior.com/cost-of-private-llm-evaluation-services/
  5. arXiv, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
  6. Acolad, "Data annotation cost". https://www.acolad.com/en/services/data-services/data-annotation-cost
  7. Zhou et al., "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  8. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data