Skip to content

Evaluation and benchmarking datasets

Real-work task benchmarks: evaluating models on professional deliverables

Quick answer

A real-work task benchmark gives a model the same brief, input files and context a professional received, then grades the work product it returns against a deliverable a client or manager actually accepted. To build one you need four linked artifacts per task: the brief, the input files, the accepted deliverable and, ideally, the reviewer feedback that got it accepted. Sourcing those from real businesses, under a license that permits evaluation use, is the hard part.

By SourceX Editorial · Updated

What separates deliverable-level tasks from QA benchmarks

Deliverable-level tasks score a finished work product, not a single answer string. A QA item asks "what is the termination notice period in this clause"; a deliverable task asks for the redline memo a partner would send to the client, with the contract, the playbook and the client's earlier emails attached. The SECQUE authors make the same point for finance: many domain benchmarks focus on isolated downstream tasks rather than the real analysis work professionals perform [1].

Occupational benchmarks of this kind, built from tasks authored by experienced practitioners and expecting documents, slides or spreadsheets as outputs, became a visible reference point for frontier labs during 2025. Once such tasks are public, they become training-contaminable, which is why frontier teams increasingly want private sets built the same way. See private evaluation sets vs public benchmarks for when held-out data is worth paying for.

Deliverable tasks also differ from agent suites. An agent evaluation task suite checks a final environment state (a ticket closed, a row written); a deliverable benchmark checks whether a 12-page variance analysis is correct, complete and usable by its reader. Many teams run both, but they need different source data.

Anatomy of a sourceable professional task

A usable task record is a bundle, not a prompt. The four components below are what to ask a supplier to locate and preserve together, with their original timestamps and file formats.

  • Brief. The instruction as the professional received it: an engagement letter excerpt, a Jira or ServiceNow ticket body, a manager's email, a statement of work line item. Keep ambiguity intact, because real briefs are underspecified and handling that is part of the skill.
  • Input files. The native artifacts: .xlsx workbooks with live formulas, .docx with tracked changes, PDFs of source contracts, CSV exports from an ERP, CAD or image files. Converting a workbook to CSV destroys the formula graph a model must reason over; see spreadsheet task benchmarks built from real workbooks.
  • Accepted deliverable. The version that was actually signed off, not a draft. Version history in SharePoint, Google Drive or a document management system such as iManage usually shows which file went to the client.
  • Reviewer feedback. Redlines, review comments, QA checklists and rework requests between draft and acceptance. This is the most valuable and least often retained component, because it encodes what an expert grader penalizes.

Practical metadata to require per task: occupation (an O*NET-SOC code is a convenient key), sector, time the professional spent, tools used, number of review rounds, and whether the deliverable was accepted as-is or after revision. Time spent lets you compare model cost and latency against human effort on the same deliverable.

How professional authorship and grading should work

Grading quality is bounded by grader expertise, so graders should come from the same occupation as the task author. LegalBench is a useful precedent: its 162 tasks across six types of legal reasoning were hand-crafted by legal professionals [2]. For deliverables, rubric grading by a single annotator is fragile; blinded comparison of a model output against the accepted human deliverable, by two or more practitioners, gives a more stable signal.

Treat the accepted deliverable as a reference, not as ground truth. Real accepted work contains errors, and audits of popular test sets have estimated average label error rates of at least 3.3% [6]. Budget for an adjudication pass in which a senior reviewer flags tasks whose accepted deliverable is itself wrong, and drop or relabel them before freezing the set.

Automated grading is attractive at scale, but calibrate it. A practical pattern is to have experts grade a stratified sample, fit the automated grader against those judgments, and report agreement by occupation, since grader drift is rarely uniform across law, accounting and engineering tasks.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy it matters
task_idfin-ap-0142Stable key across releases
occupation13-2011 Accountants and AuditorsRoutes to the right grader pool
brief_text"Reconcile Q3 vendor accruals and flag anything over $25k that lacks a PO"Preserves real ambiguity
input_filesaccruals_q3.xlsx, po_export.csv, vendor_master.csvNative formats, formulas intact
accepted_deliverableq3_accrual_recon_final.xlsx + memo.docxReference for comparison
review_rounds2Proxy for task difficulty
reviewer_feedback4 comments (missing FX adjustment; duplicate vendor)Seeds the grading rubric
human_minutes210Cost and speed comparison
pii_treatmentvendor names pseudonymized; method loggedSupports diligence
license_scopeevaluation only; no training; publish aggregate scoresGoverns reuse

Where real professional tasks come from

Most deliverable tasks live inside ordinary business systems rather than in purpose-built datasets. Good sources by function:

  • Finance and accounting: close checklists, reconciliation workbooks, audit request lists and the PBC (prepared by client) files that answered them.
  • Legal and contracts: matter intake notes, first-draft and executed agreements, negotiation redlines; see contract review evaluation datasets and the SourceX legal AI data page.
  • Engineering: design reviews, incident postmortems, change requests with linked pull requests. SWE-bench shows how far real issue-and-patch pairs can go, with 2,294 problems from 12 repositories and test-based checking [4].
  • Operations and support: escalation write-ups, SOP revisions, quality-review scorecards; these are workflow data with a natural accepted state.
  • Spreadsheet-heavy analysis: SpreadsheetBench drew 912 questions from real user problems on Excel forums, in contrast to synthesized queries and simplified files [3].

Public benchmarks age quickly. In February 2026, OpenAI stopped reporting SWE-bench Verified, stating that it had become increasingly contaminated and that score gains increasingly reflected training-time exposure [5]. Licensed tasks from private business records have never been on the open web, which is the main reason to source them; contamination-resistant evaluation design covers how to keep them that way after delivery.

Client confidentiality and rights in deliverables

Professional deliverables are usually the property of, or confidential to, a client of the supplier, so the supplier alone often cannot license them. Before scoping, ask three questions of any supplier: who owns the work product under the engagement terms, whether client confidentiality clauses or NDAs restrict reuse, and whether the inputs contain third-party data such as a counterparty's financials.

Workable answers tend to fall into a few patterns. Internal deliverables (a company's own close process or postmortems) carry fewer third-party constraints. Client deliverables may need client consent, aggregation, or heavy redaction of client identity and figures. Some material, such as privileged legal advice, may be unusable in any form. Redaction must preserve task solvability: replacing every number in a reconciliation breaks the task, so pseudonymize identifiers and keep internally consistent values.

Document the result per dataset. Data Cards offer a workable template, recording upstream sources, collection methods, intended use and decisions that affect model performance [7]. For publication rules, see publishing benchmark results on licensed evaluation data, and if the data cannot leave the supplier, consider supplier-hosted and enclave evaluation.

A buyer's request template for deliverable tasks

A precise request describes the work, not the companies that might hold it. Use this structure when briefing a supplier or a sourcing partner.

Illustrative example: invented to show structure; it does not describe an available dataset.

  1. Occupations and sectors: e.g., staff accountants, paralegals, mechanical design engineers; work performed in US business settings.
  2. Task definition: a brief plus native inputs plus an accepted deliverable; reviewer feedback preferred.
  3. Volume and spread: target task count per occupation, minimum review rounds, mix of difficulty.
  4. Formats: native files (.xlsx with formulas, .docx with tracked changes, PDF), with a manifest in JSON Lines.
  5. Exclusions: no tasks whose deliverable was never accepted; no work later found defective unless flagged.
  6. Privacy treatment: personal and client identifiers removed or replaced, method recorded, sample checked.
  7. Allowed uses: evaluation only or evaluation plus training; whether aggregate scores may be published.
  8. Grading support: whether the supplier can provide practitioner graders or rubric input.

Failure modes to design out

Most broken deliverable benchmarks fail in predictable ways:

  • Lost context. The brief references a call or a prior email thread that was not captured, so no model could reproduce the accepted answer. Have a practitioner confirm every task is solvable from its bundle.
  • Draft confusion. The "final" file is a late draft. Require evidence of acceptance: a send event, a signature, a status field in the ticketing system.
  • Format flattening. PDFs of spreadsheets, or workbooks exported to CSV, hide the structure a professional used.
  • Grader mismatch. Generalist annotators grade specialist work, rewarding fluent but wrong deliverables.
  • Over-redaction. Privacy treatment removes the facts the task depends on.

The cluster hub on LLM evaluation datasets maps the adjacent options, including outcome-labeled evaluation data when the right answer is a business outcome rather than a document, and commissioning a custom evaluation dataset when no existing records fit. SourceX's overview of evaluation sets built from real business work describes the data types involved, and you can describe the deliverable tasks you need directly.

Source professional deliverable tasks for your benchmark

SourceX sources operational datasets, including support and sales histories, engineering records, documents, and finance and legal workflows, from US companies on request; it holds no stock, and a request does not guarantee a match. Every release is approved by the supplying company, rights-reviewed, stripped or pseudonymized of personal details with the method recorded, and delivered under a license defining records, uses, term and delivery. Describe the professional tasks you need at sourcex.si/buyers.

Sources

  1. arXiv:2504.04596, "SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities" (2025). https://arxiv.org/pdf/2504.04596
  2. Hazy Research, Stanford University, "LegalBench: A collaboratively built large language model benchmark for legal reasoning". https://hazyresearch.stanford.edu/legalbench
  3. arXiv:2406.14991 (NeurIPS 2024), "SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation" (2024). https://arxiv.org/html/2406.14991v2
  4. Jimenez et al., arXiv:2310.06770 (ICLR 2024), "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770
  5. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  6. Northcutt, Athalye, Mueller, arXiv:2103.14749 (NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  7. Google Research, arXiv:2204.01075 (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data