Skip to content

Document AI data

Financial Statement Spreading Data: Balance Sheets and Income Statements from Real PDFs

Quick answer

A useful financial statement extraction dataset pairs real borrower statements (audited, reviewed, compiled and internally prepared, scanned and born-digital) with the analyst's spread: each source line mapped to a line on a standard spread template, with period, scale and sign recorded. Public research sets mostly use listed-company filings or synthetic tables, so private-company statements with spreads usually have to be licensed from businesses that hold them. Inline XBRL filings are a free, useful complement, not a substitute.

By SourceX Editorial · Updated

Why public financial statement datasets do not train spreading models

Public financial statement datasets cover listed companies or generated tables, which leaves out the messy private-company statements that commercial lenders actually spread. SynFinTabs offers roughly 100,000 synthetic financial tables, generated rather than drawn from real statements [1]. FINE extracts KPIs from SEC EDGAR filings of 18 companies as (company, time, keyword, value) tuples [2], and FinAR-Bench pairs PDFs of 100 Shanghai-listed companies with XBRL-derived tables [3].

None of these carry the label a lending model needs: the mapping from "Due from shareholder" or "Accrued liabilities and other" to a specific line in a bank's spread. Private statements also vary in ways filings do not. They include CPA compilations with "substantially all disclosures omitted," QuickBooks profit-and-loss exports, multi-column comparative periods, faxed or phone-photographed pages, and handwritten adjustments.

What a spread label contains and why it is the core asset

The analyst's spread is the most valuable label because it encodes extraction, classification and judgment in one record. A raw transcription tells a model what the page says. A spread tells it how a credit analyst reclassified each line, which items were netted, how officer compensation or related-party receivables were treated, and which total ties.

Ask for the spread together with the source document, never the spread alone. Without the PDF and page coordinates, the mapping cannot train a document model. Without the mapping, you have an OCR set, which is covered better by OCR ground truth data from real business scans and table structure recognition data.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "doc_id": "fs-000418",
  "statement_type": "balance_sheet",
  "report_level": "reviewed",
  "capture": "scanned_300dpi",
  "period_end": "2024-12-31",
  "columns": ["2024-12-31", "2023-12-31"],
  "scale": "units",
  "currency": "USD",
  "entity_masked": true,
  "lines": [
    {
      "source_label": "Due from shareholder",
      "page": 3,
      "bbox": [112, 640, 410, 658],
      "values": [184500, 92000],
      "spread_line": "NONCURRENT_RELATED_PARTY_RECEIVABLE",
      "analyst_note": "Moved out of current assets per credit policy",
      "sign": "positive"
    }
  ],
  "tie_checks": {"total_assets_equals_liabilities_plus_equity": true}
}

Coverage to specify in a financial statement extraction request

Define coverage by report level, capture condition, period structure and statement components, because each one changes model behavior. A request that only says "financial statements" will return the easiest documents to supply, which is rarely the distribution your production queue sees.

DimensionWhat to requestFailure mode if missing
Report levelAudited, reviewed, compiled and internally prepared statements, with counts per levelModel learns audited layouts and fails on compilations and QuickBooks exports
CaptureBorn-digital PDF, scanned, faxed, phone photoOCR errors on scanned pages go untested; see degraded document images
PeriodsMulti-column comparatives, interim and year-to-date, fiscal years that are not calendar yearsValues assigned to the wrong period column
ComponentsBalance sheet, income statement, cash flow, notes, supplementary schedulesDebt maturities and lease obligations in notes never extracted
Tax returnsForms 1120, 1120-S and 1065 with Schedule L and M-1, if you spread from returnsModel cannot reconcile book and tax figures
LabelsSpread mapping, analyst notes, template version, reviewer IDNo way to measure or resolve label disagreement

If you spread rent rolls or operating statements, scope those separately with rent roll and operating statement data. Bank statements used for cash-flow verification belong in bank statement and income document data.

How accounting identities validate extracted values

Financial statements carry built-in checksums, so you can validate extraction and labels automatically before any human review. Total assets must equal total liabilities plus equity, subtotals must sum their lines, and net income should roll into retained earnings once you allow for distributions.

Run these checks on both the delivered ground truth and your model output. A ground-truth row that fails its own tie-out usually means a scale error (thousands versus units), a sign error on parentheses, or a value assigned to the wrong column. These are also common error classes in tagged public filings. Flag failures for re-labeling rather than silently dropping them, because the hard cases are the ones you need.

For mapping consistency, have two analysts spread a sample of the same statements and measure agreement on spread lines with Krippendorff's alpha [5]. Low agreement on lines such as "other current liabilities" is a template ambiguity, not annotator noise, and should be fixed in the template before you scale labeling.

Using inline XBRL filings as a free complement

Inline XBRL gives you machine-readable facts aligned to their position in a human-readable filing, which makes public 10-K and 10-Q statements a cheap source of extraction ground truth. Because the tags sit inside the HTML filing itself, each tagged value can be aligned to its location on the rendered page. FinAR-Bench shows the same pairing pattern, matching original PDFs with XBRL-derived tables [3].

Know the limits before you rely on it. Filings are public-company, born-digital and tagged against the US GAAP taxonomy, which does not match a lender's spread template. Filers also use company-specific extension elements and may tag the same concept differently across periods, so normalize extensions before training. Treat XBRL as pretraining and evaluation data for table extraction, then fine-tune on private statements with spreads.

If you collect filings at scale, follow the SEC's published EDGAR fair-access rules on request rates and declared User-Agent headers, and check the current limits before building a crawler. Filings as text corpora are covered in SEC filings text corpora for financial LLMs.

Confidentiality, masking and documentation for borrower statements

Borrower identity and figures are confidential, so masking must remove identity while keeping the numeric structure the model learns from. Replace company names, addresses, EINs, owner names and bank account numbers with consistent placeholders. Keep line labels, column layout and the arithmetic relationships between values intact so that tie-out checks still pass.

Be careful with scaling or perturbing numbers to hide a borrower. It can break totals and teach the model false relationships. Personal guarantor returns such as Form 1040 carry individual data and need stricter handling than entity statements.

Ask for a dataset card covering source mix, report-level distribution, template version, annotation method and intended use [4]. When SourceX sources this kind of data, personal details such as names, emails, phones and account numbers are removed or replaced before delivery. The method is recorded and a sample is checked, though no method is perfect. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery.

How SourceX approaches financial statement requests

SourceX sources operational datasets, including finance and legal workflows and documents, from US companies on request. It also manages the commercial process, including the license and any ongoing purchases. It does not hold this data in stock, and a request does not guarantee a match. You describe the data you need, not the businesses, and every release is approved by the supplying company.

The process runs Find, Assess (the data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage. Nothing is contracted until a supplier agrees, and delivery runs through private, access-controlled workflows only after an executed agreement. If you are scoping a spreading dataset, you can describe your statement mix and spread template to SourceX. For native workbooks and models rather than statement documents, see spreadsheet and financial model datasets. For agent use cases, see finance and accounting AI training data and buyers in accounting.

Source financial statement spreading data

SourceX sources finance documents and workflow records from US companies on request, under a license that defines records, uses, term and delivery. Each release is approved by the supplying company and rights-reviewed before delivery. Request financial statement spreading data.

Frequently asked questions

Can synthetic financial tables replace real statements?

They help with layout variety and pretraining, but SynFinTabs is built from about 100,000 generated tables rather than real statements [1]. Hold out real private statements for evaluation. The tradeoffs are covered in synthetic vs real documents for document AI.

Should spreads be mapped to our template or the supplier's?

Request the supplier's original template and version, then map it to yours yourself. That keeps provenance clear and lets you re-map if your template changes.

How do I evaluate a spreading model?

Score field-level accuracy per spread line, per-period column assignment and tie-out pass rate separately. Field ground truth design is covered in document extraction evaluation sets. Browse adjacent tasks in the document AI data hub or the AI data overview.

Sources

  1. arXiv, "SynFinTabs: A Dataset of Synthetic Financial Tables for Information and Table Extraction" (2024). https://arxiv.org/pdf/2412.04262
  2. arXiv, "Enabling and Analyzing How to Efficiently Extract Information from Hybrid Long Documents with LLMs" (2023). https://arxiv.org/pdf/2305.16344
  3. Hugging Face, "FinAR-Bench dataset (Hugging Face)". https://huggingface.co/datasets/sw4tanonymous/FinAR-Bench/commit/3ddf2eed562d51004e8bc3e7d483e621bb282957
  4. Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  5. University of Pennsylvania, Annenberg School for Communication, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data