Agent, workflow and domain-reasoning data
Spreadsheet task trajectories: edit histories for spreadsheet agents
Quick answer
Spreadsheet agent training data is most useful when it shows the path, not just the finished workbook: the request that started the work, a sequence of version snapshots, cell-level diffs of values, formulas, formats and sheets between them, and the reviewed final version. Public data mostly offers forum-derived start and goal file pairs. Real edit histories come from cloud version history, saved file versions and change-tracking add-ins inside companies, and they need entity scrubbing, amount handling and formula-aware grading before you use them for SFT or evaluation.
By SourceX Editorial · Updated
Why finished workbooks are not enough for spreadsheet agents
A finished workbook tells a model what good output looks like, but not which edits produced it, in what order, or in response to what instruction. Spreadsheet agents act step by step: select a range, insert a column, write a SUMIFS, fill it down, fix a broken reference, reformat a header. Supervised fine-tuning on those steps, or rewarding an agent in a gym, needs intermediate states and the instruction that defines success.
Public research shows both the demand and the gap. SpreadsheetBench built 912 instructions from real questions on online Excel forums, deliberately moving away from synthesized queries and simplified files [1]. Spreadsheet-RL mined 18,855 ExcelForum threads with 32,691 attachments into paired start and goal spreadsheets, and reported SpreadsheetBench Pass@1 rising from 12.0% to 23.4% [2]; its released model card describes a final split of 5,928 filtered tasks trained in a multi-turn Excel gym with recalculation-based rewards [3].
Those pairs are endpoints. They do not show how an analyst rebuilt a three-statement model over a week, or how a controller rolled a forecast forward after a reviewer's comment. If you want finished workbooks themselves, the owner page for spreadsheet and financial model datasets covers that object; this page is about the trajectory.
Where real spreadsheet edit histories come from
Edit histories live in the systems companies already use to store workbooks, and their granularity varies widely by source. Treat the source system as the first spec field, because it determines what a "step" can mean.
- Cloud version history. Workbooks stored in SharePoint or OneDrive keep prior file versions, and Google Sheets keeps a revision history. Snapshot frequency depends on autosave, editing sessions and retention settings, so one "version" may contain one edit or an afternoon of work.
- Saved file versions over time. Month-end close folders often hold
Forecast_v3.xlsx,Forecast_v4_final.xlsxandForecast_v4_final_JS.xlsx. These are coarse, but they often align with real review checkpoints. - Change-tracking add-ins and audit logs. Some finance teams run add-ins or model-governance tools that log cell changes with user and timestamp. These give near-action-level granularity but cover fewer workbooks.
- Document management and ticket systems. The workbook may be attached to a request ticket, a close checklist task or an email thread, which is where the instruction usually lives.
- Commissioned recordings. When no history exists, analysts can be recorded doing defined tasks; see turning screen recordings into action-labeled trajectories and task mining data as an agent source for that route.
The trade-off is coverage versus resolution. Version snapshots cover many real workbooks at coarse resolution; add-in logs and recordings give fine resolution on fewer tasks. The cluster hub on AI agent training data compares these record types across workflows.
Turning version snapshots into cell-level action sequences
Diffing consecutive snapshots yields cell-level edits that can be ordered into an action sequence, which is the core transformation buyers should specify. The output is a list of typed operations per step, each tied to a sheet, an address and before and after content.
A workable diff pipeline parses each .xlsx (Office Open XML) version with a library such as openpyxl, reads both stored formulas and cached values, and compares them sheet by sheet. Useful operation types include value change, formula change, format change, row or column insert and delete, sheet add, rename or delete, named-range change, and data validation or conditional formatting change. Row and column inserts are the main failure mode: a naive cell-address diff reports thousands of "changes" when one row was inserted, so require alignment that detects structural shifts before emitting cell edits.
Ordering within a snapshot is not observed, so it must be inferred or left unordered. A defensible convention is dependency order: edits to precedent cells before formulas that reference them, structural edits before content edits. Label inferred ordering as inferred in the record, so you can exclude it from training on step order while still using the state transitions.
Also capture recalculation state. A cached value that does not match a recalculated formula signals manual calculation mode, an external link or a volatile function such as TODAY() or OFFSET, all of which change how you grade later.
Recovering the task instruction behind each edit session
The instruction is usually outside the workbook, so pair each history with the request that triggered the work. Without it you have state changes but no task, which is useless for instruction-following SFT and weak for evaluation.
Typical instruction sources are an email asking for a sensitivity on the discount rate, a ticket requesting a new cost center in the allocation model, a close checklist item, or a reviewer's cell comment ("tie this to the GL"). Match them to version windows by timestamp, author and file link. Where one email triggers several sessions, keep the session boundaries; ticket histories reconstructed as agent trajectories uses the same linking logic, and cross-system workflow records covers joining email, ERP and file systems into one timeline.
Ask for the instruction verbatim (after scrubbing) plus a normalized restatement, and record who wrote each. A reviewer comment that rejected a version is especially valuable, because it provides a negative example and the correction in the next snapshot.
Illustrative trajectory record for spreadsheet agent SFT and evaluation
A trajectory record should hold the instruction, ordered steps with typed operations, the snapshot references and the outcome label in one object. The schema below is a starting point for a data spec.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"trajectory_id": "wb-0193-sess-04",
"source_system": "cloud_version_history",
"workbook_type": "13-week cash forecast",
"instruction": {
"text": "Add a downside case with collections slipping 10 days and show the minimum cash week.",
"origin": "email",
"scrubbed": true
},
"start_snapshot": "snap_011.xlsx",
"end_snapshot": "snap_014.xlsx",
"reviewed_final": "snap_016.xlsx",
"steps": [
{"seq": 1, "op": "sheet_add", "sheet": "Downside", "order": "inferred"},
{"seq": 2, "op": "formula_set", "sheet": "Downside", "range": "C12:P12",
"before": null, "after": "=OFFSET(Base!C12,0,-2)*Inputs!$D$4"},
{"seq": 3, "op": "formula_set", "sheet": "Summary", "cell": "F8",
"before": null, "after": "=MIN(Downside!C40:P40)"}
],
"amount_handling": "scaled_by_constant_per_workbook",
"external_links": "broken_and_logged",
"outcome": {"label": "accepted_after_revision", "reviewer_comment_ref": "c-77"}
}
Keep the start, end and reviewed-final snapshots as files, not only as diffs, so you can rebuild state for a gym. For outcome definitions, see task success labels from business records.
Scrubbing workbooks without breaking the formulas
Spreadsheets carry identifying and confidential data in more places than cells, so scrubbing has to cover the whole package while keeping formulas valid. Entity names, customer and vendor names, account numbers and employee names appear in cell values, sheet names, defined names, comments, headers and footers, document properties (author, last modified by) and pivot caches.
Agree up front on these points:
| Item | What to agree | Common failure |
|---|---|---|
| Entity and person names | Consistent pseudonyms across all versions of a workbook | Different pseudonym per snapshot breaks diffs |
| Account and ID numbers | Replace with format-preserving tokens | Lookups (XLOOKUP, VLOOKUP) stop matching |
| Amounts | Scale or perturb, and say which, per workbook | Perturbing inputs but not hardcoded subtotals breaks ties |
| External links | Break, log the original target type | Cached values silently go stale |
| Pivot caches and hidden sheets | Scrub or drop explicitly | Raw data survives in a hidden cache |
| Metadata and comments | Strip authors, keep comment text scrubbed | Reviewer names leak via threaded comments |
Scaling amounts by one constant per workbook preserves ratios and checks such as balance-sheet ties; random perturbation does not. Decide based on whether your evaluation grades numeric values. For broader redaction practice on interaction data, see PII in screen recordings and computer-use trajectories.
Grading agent output against the reviewed final version
Evaluate agent output against the reviewed final version, and check formulas as well as values. A value-only check passes an agent that hardcoded the right number, which is exactly the habit a finance reviewer would reject.
A practical grader recalculates the agent's workbook in a headless engine, compares target ranges to the reviewed final within a tolerance, then inspects formulas for structural equivalence (references to the right input cells, no hardcoded constants where the reference used a link). Spreadsheet-RL's recalculation-based reward is one example of execution-grounded grading [3], and OSWorld's execution-based evaluation of real desktop applications shows the same principle for broader computer-use tasks [5]. For suite design, see agent evaluation task suites with state-based grading and spreadsheet task benchmarks built from real workbooks.
Two further checks matter in production. Navigation cost is real: an industry case on a spreadsheet agent at Ramp describes building a dedicated retrieval subagent because navigating workbooks and retrieving the right cells consumed a meaningful share of tool calls [4]. Hold out whole workbooks, not individual steps, so near-duplicate monthly versions of the same model do not leak between train and test.
What to ask a supplier before licensing edit histories
Ask for a short, per-dataset description of source, preparation and allowed use before any sample review. Data Cards offer a useful structure for this: upstream sources, collection method, intended use and decisions that affect model performance [6]. If you develop a generative AI system or service made available to Californians, AB 2013 requires posting documentation about the data used to train it (first due 1 January 2026, as of October 2026), so provenance fields you collect now feed that disclosure [7].
Request checklist for a spreadsheet edit-history dataset:
- Source system per workbook and the snapshot cadence it implies
- Workbook types (forecast, three-statement model, reconciliation, allocation) and counts per type
- Whether instructions are linked, and how (email, ticket, comment)
- Diff method, operation types emitted, and how row and column shifts are handled
- Scrubbing method across cells, names, metadata, comments and pivot caches
- Amount handling (scaling or perturbation) and whether ties survive
- External links and macros: removed, broken or retained
- Reviewed-final availability and how acceptance was determined
- Allowed uses (SFT, RL, evaluation) and term, written into the license
For formula-level corpora without histories, compare spreadsheet formula training data. For choosing between licensed histories and commissioned demonstrations, read licensed, commissioned or synthetic trajectories. Finance-specific agent data needs are covered on finance and accounting AI training data and licensing spreadsheets and financial models.
How SourceX handles requests for spreadsheet edit histories
SourceX sources operational datasets, including documents and finance workflows, from US companies on request; nothing is held in stock and a request does not guarantee a match. Buyers describe the data they need, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details such as names, emails and account numbers are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows after an executed agreement. You can describe your spreadsheet trajectory requirements to SourceX using the checklist above.
Source spreadsheet agent training data with SourceX
If your team needs real edit histories rather than finished files, describe the workbook types, instruction sources and grading needs you have in mind. SourceX looks for US businesses that hold that data, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything transacts. Request spreadsheet edit-history data.
Sources
- arXiv (Renmin University of China, Tsinghua University, Zhipu.AI; NeurIPS 2024), "SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation" (2024). https://arxiv.org/html/2406.14991v2
- arXiv (UIUC and Meta), "Spreadsheet-RL (arXiv:2605.22642)" (2026). https://arxiv.org/pdf/2605.22642
- Hugging Face, "Spreadsheet-RL-8B model card" (2026). https://huggingface.co/Spreadsheet-RL/Spreadsheet-RL-8B
- ZenML LLMOps Database, "Specialized retrieval subagent with reinforcement learning post-training for spreadsheet navigation". https://www.zenml.io/llmops-database/specialized-retrieval-subagent-with-reinforcement-learning-post-training-for-spreadsheet-navigation
- arXiv (XLANG Lab, HKU and collaborators; NeurIPS 2024), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
- Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.