Tables, time series and transactional data
Spreadsheet Task Benchmarks Built from Real Workbooks
Quick answer
A useful spreadsheet benchmark for LLMs pairs a real workbook with a natural-language instruction and several input/output test cases that share structure but differ in data, so a solution must generalize rather than hard-code one answer. Public sets such as SpreadsheetBench show the format works, but forum-sourced tasks are now visible to training crawlers. Teams evaluating Excel agents or spreadsheet copilots increasingly need a private, held-out set drawn from enterprise workbooks, graded on cell state, and licensed for evaluation only.
By SourceX Editorial · Updated
What SpreadsheetBench established, and where it stops
SpreadsheetBench set the reference design: 912 instructions collected from online Excel forums, backed by 2,729 test cases, with multiple test-case workbooks per instruction [1]. Its tables are deliberately messy, extending past 100 columns and 20,000 rows, with 35.7% of sheets holding multiple tables and 42.7% using non-standard relational layouts [1]. The NeurIPS 2024 authors report a substantial gap between state-of-the-art models and human performance on these tasks [2], which is the core argument for funding better eval data.
The limits matter for a commercial eval lead. Forum questions tend to be self-contained problems posted by individuals, while enterprise work involves cross-sheet models, linked workbooks, named ranges, structured references and finance conventions such as sign flips, period columns and subtotal rows. That is likely the gap a private set should fill, and it is a hypothesis to test against your own product logs rather than a settled finding.
Contamination is the second limit. Public forums are now a known training source: the Spreadsheet-RL work mined 18,855 ExcelForum threads with 32,691 spreadsheet attachments posted after 1 January 2024 [3]. LiveBench documents the general failure mode, where test data leaks into newer models' training sets and a benchmark becomes obsolete [4]. For broader context, see private evaluation sets vs public benchmarks.
Anatomy of a strong spreadsheet benchmark item
A strong item is a self-contained bundle: instruction, input workbook, answer location, several test-case variants and a grading rule. The variants are what separate a benchmark from a demo, because a model that writes =SUM(B2:B41) against one file can pass once and fail when row counts change [1].
Each item should carry:
- Instruction text as the user would write it, including ambiguity that a real analyst resolves from context (for example "last quarter" in a fiscal-year workbook).
- Input workbook in native .xlsx (OOXML) or .xlsm, preserving formulas, defined names, data validation, conditional formatting, pivot caches and hidden sheets. Converting to CSV destroys most of the signal.
- Answer position: a target range, a new sheet, or a modified table, so the grader knows what to compare.
- Test-case variants: three or more workbooks with the same schema and different values, row counts and edge cases (blanks, text-formatted numbers, merged headers).
- Expected output state per variant, computed and checked by a person who did the task.
- Task tags: cell-level vs sheet-level, formula vs restructuring vs formatting, single vs cross-sheet, and whether a pivot, lookup or array formula is required.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"task_id": "fin-recon-0147",
"instruction": "On 'Summary', fill C5:C16 with each month's unreconciled amount from 'GL_Export' minus matched items on 'Bank'.",
"modality_allowed": ["native_formula", "python_openpyxl"],
"answer_position": "Summary!C5:C16",
"tags": ["cross_sheet", "lookup", "fiscal_calendar", "sign_convention"],
"test_cases": [
{"input": "tc1_input.xlsx", "expected": "tc1_expected.xlsx", "rows_gl": 4812},
{"input": "tc2_input.xlsx", "expected": "tc2_expected.xlsx", "rows_gl": 21307},
{"input": "tc3_input.xlsx", "expected": "tc3_expected.xlsx", "rows_gl": 96, "edge": "blank_months"}
],
"grader": {"compare": "values", "tolerance_abs": 0.005, "check_formulas_present": true},
"provenance": {"source_type": "enterprise_workbook", "deidentified": true, "human_solve_minutes": 14}
}
Choosing the solution modality before you collect tasks
Decide how the agent is allowed to act before you write instructions, because the modality changes what counts as correct. SpreadsheetBench asks models to produce code that manipulates the file, and the result is checked against expected workbooks [1]. A copilot that writes native formulas, or an agent that clicks through the Excel or LibreOffice UI, needs different ground truth.
| Modality | What the grader checks | Typical failure it catches | Data you need from the source workbook |
|---|---|---|---|
| Code (Python with openpyxl or pandas) | Output cell values after execution | Hard-coded ranges, dtype coercion, lost formatting | Clean input/expected pairs per variant |
| Native formulas | Values plus presence and shape of formulas | Pasted values instead of live formulas, volatile functions, broken references | Original formula layer and dependency structure |
| UI actions (desktop agent) | Final application state | Wrong sheet focus, unsaved changes, dialog handling | Task setup scripts and state checkers |
UI-driven grading follows the execution-based pattern used by OSWorld, which sets up a real application state and checks the result after the agent acts [5]. If you grade formulas, link this work to spreadsheet formula training data, and for layout-detection subtasks see messy multi-table spreadsheets.
Realism markers to require from workbook sources
Require evidence that the workbooks look like production files, not cleaned teaching examples. SpreadsheetBench's own markers are a good floor: wide tables, tens of thousands of rows, several tables per sheet and non-relational layouts [1]. Enterprise sets should add the structures forum tasks rarely include.
Ask suppliers to report, per workbook:
- Sheet count, maximum used range, and count of formulas, defined names and external links.
- Presence of pivot tables, Power Query connections, data tables, array or dynamic-array formulas, and VBA modules in .xlsm files.
- Header pattern: multi-row headers, merged cells, period columns, subtotal and total rows.
- Workbook role: operating model, reconciliation, budget, commission calculation, inventory tracker, project plan.
- Known errors already present (#REF!, circular references), which are useful as robustness tasks if labeled.
Realism without a schema description is hard to evaluate. Pair the files with column-level documentation where it exists; see data dictionaries and schema documentation.
Building trustworthy ground truth and a human baseline
Ground truth is the most expensive and most error-prone part of the set, so budget for two independent solves per item and adjudicate disagreements. Northcutt and colleagues estimated an average label error rate of at least 3.3% across ten widely used test sets [6], and spreadsheet answers have their own traps: floating-point tails, date serials, locale decimal separators and text-stored numbers.
Practical controls:
- Solve every variant, not just the first; regenerate expected files from the adjudicated method.
- Store expected values and the solving formula or script, so you can tell "right answer, wrong method" from a true error.
- Set numeric tolerance and comparison type (value, formula, format) per item in the grader config.
- Record human solve time and success rate under the same tools the agent gets. METR's HCAST shows the pattern: baseliners with relevant professional experience work in the same environment as agents [7].
- Hold back a sealed slice for release decisions and rotate it, following the refresh logic LiveBench uses against contamination [4].
The human-vs-model gap is the number executives ask for, and it is only credible if humans worked under matched conditions [2][7].
Licensing and de-identification for a private eval set
Treat a spreadsheet eval set as confidential business data: negotiate an evaluation-scoped license, de-identify figures and names, and prohibit public release of items. Workbooks carry customer names, employee compensation, account numbers and supplier pricing in places that are easy to miss, including comments, hidden sheets, named-range labels, pivot caches and document properties.
De-identification must preserve the structure the task tests. Replacing names with consistent pseudonyms keeps lookups working; scaling all monetary values by one factor keeps totals reconcilable; shifting dates must respect fiscal calendars. If any workbook holds protected health information from a HIPAA covered entity or business associate, de-identify it under the Safe Harbor or Expert Determination method [8]. A practical method write-up is at how to de-identify spreadsheets and financial models.
On the license itself, confirm allowed uses (internal evaluation, regression testing, publishing aggregate scores), whether items may be shown to annotators or vendors, retention, and whether results can be published. See publishing results on licensed eval data and the spreadsheet and financial model licensing guide.
Request checklist for a spreadsheet agent evaluation set
A precise request lets suppliers judge fit quickly and avoids paying for workbooks you cannot grade. Use this as a starting brief.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example entry |
|---|---|
| Target product | Excel add-in copilot writing native formulas; Python fallback allowed |
| Workbook types | Monthly close reconciliations, sales commission models, inventory trackers |
| Format | Native .xlsx/.xlsm with formulas, names and pivots intact |
| Task count and split | 300 instructions; 3+ test-case variants each; 20% sealed |
| Task mix | 40% cross-sheet lookups, 25% restructuring, 20% aggregation, 15% error repair |
| Ground truth | Double-solved, adjudicated expected workbooks plus solving method |
| Human baseline | Solve time and success under matched tools |
| Privacy | Names, account numbers and compensation replaced; method recorded |
| License | Evaluation and regression testing only; no redistribution |
| Exclusions | No forum-posted or public-template workbooks |
Related deliverables have their own guides: private text-to-SQL evaluation sets, agent evaluation task suites and real-work task benchmarks on professional deliverables. The structured data buyer's guide maps the full cluster.
How SourceX fits a spreadsheet benchmark request
SourceX sources operational datasets, including documents and finance and legal workflows, from US companies, and manages the licensing process. Data is sourced on request rather than held in stock, so a request does not guarantee a match; every release is approved by the supplying company. Datasets are rights-reviewed and delivered under a license defining records, uses, term and delivery, with personal details removed or replaced before delivery and the method recorded. You can describe the workbook types and task design you need on the buyer intake page. For background, see spreadsheet and financial model datasets, do AI labs buy spreadsheets? and evaluation datasets from real business work.
Source a private spreadsheet benchmark from real workbooks
SourceX looks for US businesses that hold the workbook data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Nothing is contracted until a supplier agrees. Describe your spreadsheet eval set on the buyers page.
Frequently asked questions
Can I just use SpreadsheetBench for release decisions?
It is a sound public baseline, but its tasks come from public forums [1], and forum content is now used to train spreadsheet agents [3]. Use it for comparability and a private set for release gates.
How many test cases per instruction are enough?
SpreadsheetBench averages about three per instruction [1]. Add variants that change row counts, introduce blanks and shift date ranges, since those expose hard-coded solutions.
Should the grader check formulas or only values?
Check values for code-modality agents. For copilots that write formulas, also check that formulas are present and reference the right ranges, because pasted values pass a value check while breaking the workbook for the user.
Sources
- arXiv (Ma et al.), "SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation" (2024). https://arxiv.org/html/2406.14991v2
- NeurIPS, "SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation (NeurIPS 2024 poster)" (2024). https://neurips.cc/virtual/2024/poster/97753
- arXiv, "Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning" (2026). https://arxiv.org/pdf/2605.22642
- arXiv (ICLR 2025), "LiveBench: A Challenging, Contamination-Free LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
- arXiv (Xie et al.), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
- arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- arXiv (METR), "HCAST: Human-Calibrated Autonomy Software Tasks" (2025). https://arxiv.org/pdf/2503.17354
- eCFR / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.