Evaluation and benchmarking datasets
Outcome-labeled evaluation data: real business decisions as ground truth
Quick answer
An evaluation dataset with real-world outcome labels pairs each historical case, rebuilt as it stood when someone decided it, with what happened next: the claim was paid, the denial was overturned on appeal, the refund held with no chargeback. Treat an outcome as ground truth only when it is final, confirmed by a later event, and judged under the policy version then in force. Remove every field written after the decision, and grade the agent's decision rather than its wording.
By SourceX Editorial · Updated
To assemble a whole golden set, see building a golden evaluation dataset from real business records; to test whether outcome fields are reliable, see verifying ground truth in operational records. The evaluation datasets hub maps the cluster.
Decisions and outcomes are two different labels
A business record usually holds a decision, made by a person under a policy at a point in time, and sometimes a later event showing whether that decision held. Scoring an agent against decisions measures how closely it imitates past staff; scoring it against verified outcomes measures whether its calls were right.
Practitioner guidance already treats human corrections, escalations and refunds as sources of eval cases [1], and a practical guide to LLM evaluation asks that eval data match the tasks the system will perform and stay distinct from training data [2]. What differs by workflow is where the confirming event lives and when it arrives.
| Workflow | Decision recorded | Later event that confirms or contradicts it | Outcome exists only when |
|---|---|---|---|
| Insurance claim | Pay, deny, partial pay; reserve | Reopen, appeal or complaint result, subrogation recovery, suit | The claim closed and appeal rights lapsed |
| Warranty claim | Approve, reject, adjust labor or parts | Dealer appeal result, returned-part inspection, charge-back to the dealer | The part was returned or the rejection appealed |
| Accounts payable exception | Approve, hold, reject invoice | Paid, voided, credit memo, duplicate-payment recovery | Payment ran and any recovery audit finished |
| Credit or limit decision | Approve, decline, counteroffer | Repayment, delinquency, charge-off | The application was approved |
The last column is the one buyers miss: each outcome is observed only on some branches of the decision (see selective labels below).
Which outcomes count as trustworthy ground truth
Trust an outcome as a label when it is final, confirmed by an event the original decision-maker did not control, and judged against the policy version in force. Grade every case by the strength of its evidence so that scores on the strongest labels can be reported on their own.
| Label grade | Evidence on the record | How to use it |
|---|---|---|
| A: verified outcome | An independent later event: appeal decided, part inspected, or payment cleared with no recovery or dispute by the end of the window | Primary scoring set |
| B: reviewed decision | Second-level sign-off, QA review or audit sample, and no contrary event within the window | Score, but report separately from grade A |
| C: unreviewed final decision | Closed by the original handler; no downstream evidence either way | Agreement metrics only; send a sample to expert re-adjudication |
| Exclude | Auto-closed for inactivity, default or system-set values, appeal window still open, case still open | Do not score |
"Resolved" and "closed" are workflow states, not outcomes: a ticket can close because the customer stopped replying. Reversals are the most informative records, because they show where staff judgment and later evidence disagreed (human override and correction logs).
Five ways outcome labels mislead an evaluation
Outcome labels fail in predictable ways: the recorded decision was wrong, the outcome was never observed, the policy changed, the outcome had not arrived yet, or the inputs leaked it. Each has a specific control; a supplier who cannot support the controls is offering decisions, not ground truth.
1. Wrong decisions recorded as truth. Even curated benchmarks carry label errors: Northcutt et al. estimated an average error rate of at least 3.3% across the test sets of 10 widely used datasets and showed that such errors can change model rankings [3]. Operational records were never produced as test labels, so have two domain reviewers re-adjudicate a random sample against the policy text and report Krippendorff's alpha, where 1 is perfect agreement and 0 is agreement no better than chance [4]. Automated screening helps prioritize, but in Northcutt et al.'s audit human reviewers confirmed only about half of the flagged candidates [3], so treat flags as a review queue (gold-label audits and adjudication).
2. Selective labels. An outcome exists only on the branch the decision allowed. Declined applicants never repay or default, dismissed fraud alerts are often never investigated, and rejections nobody appealed get no second look. When an agent approves a case a human declined, there is no outcome to score it against. Report that share as unscorable, never as an error or a pass, and send a sample to expert reviewers.
3. Policy drift. A changed refund window, approval limit, coverage rule or code set turns yesterday's correct decision into today's wrong one. Every case needs the policy version in force when it was decided, and cases keyed to a retired version must be relabeled or dropped; see code-set revisions and process changes in multi-year datasets.
4. Outcome lag. Recent cases have not had time to reopen, be appealed, charge back or default, so they look cleaner than they are. Set an observation window per outcome type and exclude cases whose window is incomplete. This pulls against contamination control, which favors recent records (post-cutoff evaluation data): the usable band is cases old enough to have matured and never published.
5. Leakage from post-decision fields. Closing notes, payment status, appeal flags, final reserves and adjuster summaries are written after the decision. If any of them reaches the agent's input, the test measures reading, not judgment. Rebuild each input as of the decision timestamp from an audit log or field history.
Historical decisions can also encode bias against groups of people, and outcome labels inherit it; see auditing historical decision bias in operational labels.
Grading an agent on decisions, not reference text
Compare the agent's structured decision (action, amount, reason code, routing) with the verified outcome, the way agent benchmarks compare end states rather than transcripts. Report each error type separately, and score the original human decisions on the same cases as a baseline.
τ-bench grades a conversation by comparing the final database state with an annotated goal state [5]. Business records supply the equivalent end state directly: the decision fields. Text similarity to the handler's note would reward an agent that writes like the handler and decides wrongly.
| Decision element | Metric | Watch for |
|---|---|---|
| Action (approve, deny, refer) | Confusion matrix; cost-weighted error | False approvals and false denials cost different amounts |
| Amount (payment, reserve, refund) | Within-tolerance rate; absolute error | Take the tolerance from policy rounding and authority rules |
| Reason or denial code | Exact match or set overlap | Code lists change between policy versions |
| Routing or escalation | Match on target queue or approver | Escalating can be correct where the human decided alone |
| Written rationale | Rubric scored by people, or by an LLM judge checked against them | See LLM-as-a-judge calibration sets |
Keeping only grade A labels shrinks the set, and slices can fall to dozens of cases. A position paper at ICML 2025 argues that CLT-based confidence intervals are too narrow for evals with fewer than a few hundred items [6]. Size slices with eval set statistical power, and oversample rare, costly outcomes with stratified evaluation sets.
Worked example: agreement with staff versus verified accuracy
Scoring the same agent against raw decisions and against verified outcomes can change both the headline number and the errors you see.
Illustrative example: invented to show structure; it does not describe an available dataset.
A team tests a warranty-claims agent on 400 closed claims, each with grade A evidence. Staff approved 300 and rejected 100. Later events show staff were wrong on 27: dealers won appeals on 18 rejections, and 9 approvals were charged back after returned-part inspection found no defect. A further 120 rejections with no appeal and no inspection were held out as an unscorable slice.
The agent disagrees with staff on 40 claims. In 14 it approves a rejection that was later overturned and in 4 it rejects an approval that was later charged back, so 18 disagreements are correct. The other 22 are agent errors: 6 rejections of valid claims and 16 approvals of invalid ones.
| Measure | Agent | Original staff decisions |
|---|---|---|
| Agreement with original decisions | 360/400 = 90.0% | 100% by definition |
| Accuracy against verified outcomes | 369/400 = 92.25% | 373/400 = 93.25% |
| Wrongful approvals (invalid claims paid) | 21 | 9 |
| Wrongful rejections (valid claims refused) | 10 | 18 |
Against raw decisions the agent looks ten points short and is penalized for 18 correct calls. Against verified outcomes it trails staff by one point, with errors running the other way: fewer valid claims refused, more invalid ones paid. Whether that trade is acceptable is a cost question a single accuracy score hides.
What to request with each outcome-labeled case
Ask for the inputs as they stood at decision time, the decision with its policy version and the decider's authority, and every downstream event with its date. You can then derive labels yourself and re-derive them when rules change; a label without its evidence cannot be audited.
Illustrative example: invented to show structure; it does not describe an available dataset.
case_id: wc_5d21e8 # pseudonymized; same key in claim, payment, appeal and parts tables
case_type: warranty_claim
decided_at: 2025-11-03T16:40:00Z
inputs: # rebuilt as of decided_at from the field-history log
repair_order: {op_code: valve_body_replace, labor_hours: 3.2, parts_amount: 412.00}
vehicle: {model_year: 2023, odometer: 38410, in_service: 2023-02-14}
technician_comment: "harsh 2-3 shift when warm; stored codes; replaced valve body"
attachments: [repair_order.pdf, scan_tool_report.pdf]
decision:
action: reject
reason_code: insufficient_diagnosis
decided_by_role: warranty_administrator
authority_limit: 1500.00
automated_rule: none
policy: {doc_id: powertrain-warranty-terms, version: 2025-07}
downstream: # never shown to the agent
appeal: {filed: 2025-11-10, decided: 2025-12-02, result: overturned, by_role: regional_warranty_manager}
payment: {paid: 2025-12-09, amount: 1036.80, recovered: false}
part_return: {inspected: 2026-01-15, finding: defect_confirmed}
observation_window_days: 180
window_complete: true
label:
action: approve
amount: 1036.80
grade: A
rule: "a decided appeal overrides the original decision; inspection must not contradict it"
slices: [powertrain, overturned_on_appeal, policy_2025-07]
Ask the supplier to document how each label was derived (which events count, which window applies, who records each event) alongside the motivation, composition, collection process and recommended uses that Datasheets for Datasets proposes documenting [7]. For domain versions, see insurance claims AI evaluation and automotive warranty claims data.
Rights, privacy and regulatory checks for decision records
Decision records usually describe people, and outcome evidence often sits in a second system or comes from a third party, so confirm rights for both halves of every label. In regulated decision areas, test data itself can fall under data-governance and documentation rules.
- Outcome evidence from other parties. Appeal results, card-network disputes, payer remittances and dealer charge-backs may come from systems or counterparties the decision-maker does not control; confirm the supplier may license those fields, not only its own decisions.
- Financial records. Regulation P (12 CFR 1016.11) limits a recipient's reuse and redisclosure of nonpublic personal information received from a nonaffiliated financial institution, even when the recipient is not a financial institution [8].
- De-identification that keeps windows intact. Observation windows, appeal deadlines and policy effective dates must survive, so use one date offset per case and consistent surrogate keys across tables (de-identifying evaluation data without breaking the test).
- EU AI Act. For high-risk systems, Article 10 requires training, validation and testing data sets to be relevant, sufficiently representative and, to the best extent possible, free of errors and complete, and requires examination of possible biases; where no model training is used, these requirements apply only to the testing data sets [9]. As of October 2026, the Act has been amended by Regulation (EU) 2026/1744 [10]; that regulation also amends Article 10 [10], and commentary reports that Annex III high-risk obligations now apply from 2 December 2027, so confirm dates and wording against the consolidated text.
- Colorado. SB26-189, signed 14 May 2026, requires developers of automated decision-making technology that materially influences consequential decisions (including employment, housing, lending, insurance and health care) to give deployers documentation of intended uses, training data categories, known limitations and guidance on human review from 1 January 2027 [11]. As of October 2026, the Attorney General's interim draft rules (released 6 October) are open for comment until 26 October [12]; see Colorado SB 26-189 training data documentation.
SourceX sources operational datasets from US companies, including finance and legal workflows and support histories, and manages the licensing process. It sources on request, not from stock, so a request does not guarantee a match; the supplying company approves every release, and each dataset goes through rights review and is delivered under a license defining the records included, allowed uses, term and delivery.
Names, account numbers and similar personal details are removed or replaced before delivery and a sample is checked, though no de-identification method is perfect. To scope a set, describe the decisions and outcome evidence your evaluation needs; record types are outlined in AI evaluation datasets built from real business work and workflow data.
Specifying an outcome-labeled eval set: checklist and warning signs
A request should name the decision, the outcome evidence, the observation window and the label rules, not only a case count. Treat each warning sign as a reason to ask more before accepting a delivery.
Specify:
- The decision under test: allowed actions, amounts and reason codes
- Outcome events that confirm or reverse a decision, and their source systems
- An observation window per outcome type, with an incomplete-window flag
- Policy documents and versions covering the whole case period
- Inputs rebuilt as of the decision timestamp, with the decider's role and authority, and the list of excluded post-decision fields
- A label grade per case; minimum counts per slice for reversals, denials and escalations
- De-identification that preserves dates, intervals and cross-table keys
- License scope: evaluation use, hosted-API exposure and publication of scores (evaluation-only data license terms)
Warning signs:
- "Resolved" or "closed" status offered as the only label
- Current-state extracts with no field history or audit log
- No reversals, appeals or reopens anywhere in the sample: they were filtered out or never captured
- Outcomes reported for declined or dismissed cases with no account of how they were observed
- Labels you are not allowed to re-derive from the underlying events
Need decisions with verified outcomes to test against?
Describe the decisions your agent makes, the outcome evidence that would confirm them, the observation window and the license scope you need. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Describe your evaluation data needs.
Sources
- OneUptime blog (practitioner guide), "How to Build a Golden Evaluation Dataset from Real LLM Production Failures" (2026). https://oneuptime.com/blog/post/2026-08-31-build-golden-evaluation-dataset-production-failures/markdown
- arXiv:2506.13023, "A Practical Guide for Evaluating LLMs and LLM-Reliant Systems" (2025). https://arxiv.org/html/2506.13023v1
- Northcutt, Athalye, Mueller (arXiv:2103.14749; NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Klaus Krippendorff (University of Pennsylvania, Annenberg School for Communication), "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
- Yao, Shinn, Razavi, Narasimhan (Sierra; arXiv:2406.12045), "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- ICML 2025 (poster), "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
- Gebru, Morgenstern, Vecchione, Wortman Vaughan, Wallach, Daumé III, Crawford (arXiv:1803.09010), "Datasheets for Datasets" (2018; CACM 2021). https://arxiv.org/pdf/1803.09010
- Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2026/1744 amending Regulations (EU) 2024/1689, (EU) 2018/1139 and (EU) 2023/1230 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
- Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
- Colorado Attorney General, "Colorado Automated Decision-Making Technology & Chatbot Safety Rulemaking" (2026). https://coag.gov/ai/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.