Skip to content

Real-world spreadsheet and financial model datasets for AI training

A spreadsheet dataset is a collection of real working files, such as budgets, forecasts, financial models and operational trackers, delivered with formulas, cross-sheet references, named ranges, pivot tables and macros intact rather than flattened to values. SourceX sources these workbooks from established businesses' Excel, SharePoint, OneDrive and Google Sheets archives, with version history where the platform kept it. Names are replaced consistently in cells, sheet names and formula text, so the models still recalculate after de-identification.

Dataset manifest

Sourced to your spec
What it is
Working Excel and Google Sheets files with formulas, links and version history
Typical systems
Excel on SharePoint or OneDrive, Google Sheets, Box, Dropbox, file shares
Typical history
Several fiscal years of models; file archives often span 3–15+ years
Modality
Native workbooks plus cell-level formulas, cached values and workbook structure
Delivery formats
Native xlsx or xlsm files plus cell-level JSONL keyed by file and cell
Preparation
Names de-identified in cells, sheet names, formulas, comments and hidden caches
Licensing
Permitted use, derivatives and deletion terms fixed per license; client models need client approval
Availability
Sourced to your spec from partner finance and operations teams; not guaranteed

What a delivery contains

Fields vary by source system and are fixed per order. A typical delivery includes:

FieldTypeWhat it holds
file_idstringPseudonymous workbook key that joins the native file, its extracted cells, its earlier saves and any linked files.
formatobjectFile type such as xlsx, xlsm, xlsb or Google Sheets, the application that last saved it, and calculation settings.
model_typeenumWhat the workbook does, such as budget, rolling forecast, three-statement model, valuation, pricing calculator, capacity plan or tracker.
sheetsarrayEvery sheet with its role (inputs, data, calculation, output, checks) and visibility, including hidden and very hidden sheets.
cellsarrayEvery non-empty cell with sheet, address, nearby row and column labels, number format and input or calculation styling.
cells[].formulastringThe formula as written, with shared formulas expanded so each cell carries its own text.
cells[].valuemixedThe value the cell last calculated to, typed as number, text, date, boolean or error.
dependenciesarrayPrecedent-to-dependent edges at cell or range level, including cross-sheet and external references.
named_itemsarrayDefined names and Excel tables with the ranges they point to, so structured references resolve.
pivotsarrayPivot table definitions with source range, row, column, value and filter fields.
codeobjectVBA modules and Power Query steps as text, plus data connections with credentials removed.
external_linksarrayLinks to other workbooks, whether the target is in the delivery, and the values last cached from it.
versionsarrayEarlier saves of the workbook, each with a cell-level diff against the one before it.
linked_outputsarrayMemos, decks or emails that used the model's figures, where they are in scope.

Example record

{
  "file_id": "wb_4e19a7",
  "path": "/Finance/FP&A/FY2024 plan/Operating model v14.xlsx",
  "format": { "type": "xlsx", "last_saved_by": "excel", "calc_mode": "automatic",
              "iterative_calc": true },
  "model_type": "three_statement_model",
  "sheets": [
    { "name": "Inputs", "role": "inputs" }, { "name": "Data", "role": "data" },
    { "name": "Rev_build", "role": "calculation" }, { "name": "IS", "role": "output" },
    { "name": "BS", "role": "output" }, { "name": "Debt", "role": "calculation" },
    { "name": "Checks", "role": "checks" }, { "name": "Customer_mix", "role": "pivot" },
    { "name": "Old_cases", "role": "archive", "visibility": "hidden" }
  ],
  "cells": [
    { "sheet": "Inputs", "ref": "C12", "label": "FY24 price increase", "value": 0.045,
      "style": { "format": "0.0%", "convention": "input" } },
    { "sheet": "Rev_build", "ref": "F18", "value": 431459.6,
      "formula": "=SUMIFS(Data!$H:$H,Data!$B:$B,\"[CUSTOMER_017]\",Data!$D:$D,F$3)*(1+Inputs!$C$12)" },
    { "sheet": "IS", "ref": "E8", "formula": "=Rev_build!F212", "value": 48211975 },
    { "sheet": "IS", "ref": "E10", "formula": "=E8*Inputs!$C$13", "value": 19863333.7 },
    { "sheet": "IS", "ref": "E31", "formula": "=AVERAGE(Debt!D22:E22)*InterestRate",
      "value": 1315875, "circular": true },
    { "sheet": "Checks", "ref": "C5", "label": "Balance check",
      "formula": "=ROUND(BS!E40-BS!E72,0)", "value": 0 }
  ],
  "named_items": [{ "name": "InterestRate", "refers_to": "Inputs!$C$41", "value": 0.0825 }],
  "pivots": [{ "sheet": "Customer_mix", "source": "Data!$A$1:$H$5121", "values": ["sum_of_revenue"] }],
  "code": { "vba_modules": [], "power_query": ["pq_gl_actuals"] },
  "external_links": [{ "target": "[EXT_WORKBOOK_1]", "in_delivery": false, "cached_values": true }],
  "versions": [
    { "v": 9, "t": "2023-10-02T16:10:44Z", "author_role": "fpa_analyst" },
    { "v": 14, "t": "2023-11-20T08:31:02Z", "author_role": "fpa_manager", "diff": [
      { "sheet": "Inputs", "ref": "C12", "from": 0.06, "to": 0.045 },
      { "sheet": "IS", "ref": "E31", "from": "=Debt!E22*InterestRate",
        "to": "=AVERAGE(Debt!D22:E22)*InterestRate" } ] }
  ],
  "deid": { "replaced_in": ["cells", "formula_text", "sheet_names", "comments", "pivot_cache"],
            "recalc_matches_original": true }
}

Synthetic record for illustration. Field names, structure and format are agreed per order.

What AI teams use it for

Train spreadsheet agents to build and extend models

Real workbooks show how analysts lay out inputs, calculations and outputs, so an agent learns to add a scenario or a forecast year that fits the structure already there.

Translate between requests and formulas

Row and column labels next to the formulas that compute them give grounded pairs for writing formulas from a plain-language request and for explaining formulas that already exist.

Detect and fix model errors

When a later save fixes a hard-coded number, a range that stops short or a wrong reference, the version pair labels a real error together with its fix.

Answer questions over complex workbooks

Multi-level headers, merged cells, subtotals and cross-sheet chains test table understanding well beyond the tidy grids found in most public table benchmarks.

Evaluate agents on held-out edits

An earlier version plus the change request behind the next one makes a graded task with a known human answer.

Use-case guides: Finance and accounting agents, Document AI, enterprise search and RAG, Enterprise and computer-use agents

What makes this data valuable

Live formulas

Formulas rather than pasted values keep the logic that connects inputs to outputs.

Dependency depth

Cross-sheet chains, named ranges and lookups are the hard part for agents, so depth makes harder training items.

Version chains

Successive saves show how a model was built, corrected and reworked for new questions.

Modeling conventions

Separate input, calculation and output sheets, input styling and check cells reflect professional practice.

Feature range

Pivots, data tables, array formulas, Power Query and macros cover more of what people actually build.

Linked outputs

Memos and decks that quote a model's figures tie each workbook to the question it was built to answer.

Why spreadsheet agents need the formulas

Spreadsheet agents need the formulas because a financial model behaves more like a program than a document. The figures on its output sheet are one run of it; the logic lives in the formulas, references and named ranges that carry assumptions through to results. Exports that flatten a workbook to values keep the run and lose the program, and most tables published on the web are already flattened. That logic is what a spreadsheet agent has to read, extend and debug.

Two of the most widely studied public corpora with formulas, the Enron spreadsheet corpus and the EUSES corpus, hold files that are more than twenty years old. They predate structured table references, Power Query, XLOOKUP, dynamic arrays and LAMBDA, and web-crawled collections tend to favor templates and published data tables over working models. Current models from operating finance teams carry the conventions an agent must respect: inputs kept apart from calculations, check cells that should read zero, scenario switches, and interest lines that are circular on purpose, with iterative calculation turned on. A request such as "add a downside case" only tests an agent when it has to fit that structure.

For evaluation, grade spreadsheet tasks by recalculating and comparing outputs rather than matching formula text, because many different formulas return the same value. Working models also carry their own verifiers: a balance check that must read zero, or a cash total that must tie across sheets, will often fail when an edit breaks the logic, which gives a second test that needs no human grader.

Where formula extraction goes wrong

Formula extraction usually goes wrong where formulas meet cached values, because an xlsx file stores each formula next to the value it last calculated. Read only the values and the logic is gone; read only the formulas and you must recalculate everything, and some functions will not evaluate outside the application that wrote them. Cached values are missing when a file was written by a script or export tool that never calculated it, and can be stale when a workbook was left in manual calculation mode.

Parsers also trip on storage shortcuts. A formula filled down a column can be stored once and shared by the cells below it, and a dynamic array formula sits in a single anchor cell while its results spill into neighbors that hold only values. Count formulas naively and you misjudge how much logic a workbook holds.

Workbooks also carry content nobody sees on screen: hidden and very hidden sheets, pivot caches that can keep a full copy of their source data, and values cached from linked external workbooks. That matters for privacy review, and for training as well, since a model that learns from those layers learns from data the workbook's users never looked at.

What to check before licensing

  • Confirm whose numbers they are. Models built by accounting, advisory or consulting firms usually hold client financials under engagement confidentiality, and the client may have to authorize licensing.
  • Ask how workbooks about listed companies or live transactions were screened, since deal models and forecasts can contain material non-public information.
  • Check de-identification beyond visible cells — hidden and very hidden sheets, comments and notes, headers and footers, defined names, document properties, VBA code, pivot caches and external-link caches.
  • Ask whether replacements were made consistently in cell values, formula text and sheet names, and whether each workbook was recalculated afterwards with outputs matching the original cached values.
  • Measure formula density on the sample, meaning the share of workbooks with live formulas, so CSV dumps and pasted-value copies do not dominate the delivery.
  • Recalculate a sample of workbooks yourself and compare the results with their cached values. Mismatches point to stale caches, functions your engine does not support, or links that no longer resolve.
  • Confirm what happened to external links — whether the linked workbooks are included, or only the values they last returned.

How licensing works through SourceX

  1. 1

    Define

    Send the domain, modality, volume, format, timeline and permitted use you need.

  2. 2

    Source

    SourceX identifies businesses that hold matching data and are open to licensing it.

  3. 3

    Qualify

    Fit, rights and quality are checked, and you review samples before committing.

  4. 4

    License

    Scope, permitted use, exclusivity, price and obligations are agreed in writing.

  5. 5

    Deliver

    Approved data is prepared, de-identified where required and transferred securely.

Questions buyers ask

Do licensed workbooks keep their formulas, or only values?

They keep their formulas when the source files contain them. Each formula cell can be delivered with the formula as written and the value it last calculated to, alongside the native file, so you can study the logic and check results without recalculating every workbook. Values-only files, such as CSV dumps and pasted-value copies, can be flagged or excluded, and you can set a minimum formula share in your request.

How are names removed from inside spreadsheets?

Names are replaced wherever they occur, not only in visible cells. Customer, employee and client names also turn up in sheet names, formula text, comments, headers and footers, defined names, document properties and pivot caches. Replacements must be consistent, because a lookup or SUMIFS that matches on a customer name breaks if the name changes in the data but not in the formula. Recalculating afterwards and comparing outputs with the original values shows whether each model still works.

Are the financial figures real?

They come from working models, which is much of their value, but figures can also identify a business. A revenue line, a headcount plan and a list of sites together may point to one company. Depending on scope, identifying labels are removed, inputs are scaled and the model recalculated so it stays internally consistent, or sensitive models are excluded. The approach is agreed for each dataset.

Are macros, Power Query steps and scripts included?

They can be. VBA modules travel inside macro-enabled files such as xlsm and xlsb, and Power Query steps are stored in the workbook, so both can be extracted as text and reviewed for credentials, file paths and server names. Office Scripts work differently: Excel saves them as separate .osts files in OneDrive or SharePoint rather than inside the workbook, so they are included only if the partner exports them as well.

Which kinds of spreadsheets can be sourced?

That depends on what partners built and agree to license: budgets and rolling forecasts, three-statement and valuation models, pricing and quoting calculators, capacity and headcount plans, commission calculators, trackers and reconciliation workbooks. Corporate finance teams hold their own plans; accounting, advisory and consulting firms hold many models built for clients, which usually need client authorization. Naming the model types, industries and level of complexity makes matching faster.

Is version history available for workbooks?

Often, in two forms: revisions kept by the storage platform, and the save-as chains analysts build themselves, such as v12, v13 and final. Either kind can be compared cell by cell, which separates changed inputs from rewritten formulas and new rows or sheets. That cell-level difference is what turns a pair of versions into a labeled task, so ask for version pairs explicitly if you plan to build edit or error-correction data.

Can Google Sheets files be included?

Yes, though the export route matters. Downloading a Google Sheet as an xlsx file converts functions that both products share, but functions that exist only in Sheets, such as IMPORTRANGE, GOOGLEFINANCE and QUERY, cannot run in Excel, so the downloaded copy keeps their last results rather than working formulas. Exporting through the Sheets API can return each cell's formula exactly as written together with its computed value, so that route is preferred.

Evaluating this data for procurement?

Diligence packets are prepared per dataset. Rights, privacy processing and quality differ between datasets.

Request dataset diligence

Tell us what your models need

Send your spec — domain, volume, format, timeline and permitted use — and SourceX will match it against partner data and come back with what can be licensed.

Updated 3 October 2026. Own data like this? See how companies license it to AI developers.

See if you qualify