Skip to content

Document AI data

Table Structure Recognition Data from Real Financial and Operational Documents

Quick answer

A table structure recognition dataset pairs table images or PDF regions with cell-level labels: row and column boundaries, spanning cells, header roles and cell text. Public sets such as PubTabNet, PubTables-1M and FinTabNet are large but come mostly from scientific articles [1] or a narrow band of company annual reports. Teams extracting invoices, statements and operational reports usually need licensed real documents with borderless, multi-level, scanned and multi-page tables labeled in a logical format like HTML or JSON.

By SourceX Editorial · Updated

What a table structure recognition label actually contains

A usable TSR label describes both the physical grid and the logical table, not just a bounding box around the table. Detection labels (one box per table, as in COCO-style JSON) answer where a table is; structure labels answer how its cells relate. Production pipelines need both, plus functional roles that tell a downstream parser which cells are column headers, row headers (PubTables-1M calls the latter projected row headers) and data.

A complete record typically carries five layers:

  • Page reference: the source page image at a stated DPI, or the PDF page with its coordinate space, plus rotation and crop.
  • Table region: a polygon or box, with a flag for tables that continue on the next page.
  • Grid: row and column separators, or row and column objects with coordinates.
  • Cells: each cell's start and end row and column (spans), bounding box, text, and header or data role.
  • Logical serialization: HTML with rowspan and colspan, as PubTabNet uses, or an equivalent JSON or markdown target for vision-language model supervised fine-tuning [1].

If any layer is missing, some model families cannot train on the data. Image-to-sequence models need the HTML target; DETR-style object detectors such as the Table Transformer need row, column and spanning-cell boxes.

Where PubTabNet, PubTables-1M and FinTabNet come from

The large public TSR sets are built by aligning a structured source with a rendered page, which is why they skew toward scientific publishing. PubTabNet aligned the XML and PDF versions of PubMed Central open-access articles to produce about 568k table images with HTML labels [1]. PubTables-1M, from Microsoft, also draws on scientific articles; its authors focused on reducing oversegmentation and inconsistent annotation at roughly a million tables.

FinTabNet came out of IBM's Global Table Extractor work, which used automatic labeling to produce cell structure annotations from annual reports of S&P 500 companies; check its current access and license terms before relying on it. TableBank used weak supervision from online documents to scale table detection and recognition labels [2]. These sets are strong for pretraining, but their layouts, fonts and capture conditions differ from a bank statement scanned at an angle or an ERP export printed to PDF.

Two recent trends confirm the gap. CISOL annotated real construction-industry documents, more than 800 images with over 120k instances, because generic sets did not cover that domain [3]. On the other side, SynFinTabs generated about 100,000 synthetic financial tables but still built a small real-world test set to check transfer, and a 2026 dataset uses a multi-agent LLM system to synthesize table images at scale [4][5]. For when synthesis holds up, see where generated documents break.

The hard cases production tables add

Real business tables fail models in a short list of repeatable ways, and a dataset should be specified against that list. Scientific-article tables are mostly born-digital, single-page and consistently typeset, so they underrepresent the cases below.

  • Borderless tables: alignment and whitespace define columns, common in financial statements and management reports. Models trained on ruled tables merge adjacent columns.
  • Multi-level and spanning headers: "Three months ended" spanning two year columns, or region headers spanning several product columns. Errors here corrupt every value beneath.
  • Projected row headers and indentation hierarchy: section labels like "Current assets" that span the full width, and indented subtotals whose hierarchy lives only in indentation.
  • Tables broken across pages: repeated or omitted headers, "continued" markers and footers in between. Labels need a cross-page table ID.
  • Scans, faxes and phone photos: skew, blur, bleed-through and stamps over cells. Public scanned-document sets are small; FUNSD, for example, has 199 forms [6]. See degraded scans, faxes and phone photos.
  • Dense numeric tables: parentheses for negatives, footnote markers, currency columns and empty cells that are meaningful (a dash versus a blank).
  • Nested and mixed content: line-item tables inside invoices with wrapped descriptions, or key-value blocks that look tabular.

When you write a request, give a target share for each case rather than a single total count. A set of 5,000 tables that is 90 percent ruled and single-page will not move borderless or multi-page accuracy.

Choosing TEDS, GriTS or cell adjacency for evaluation

The right metric depends on what your downstream system consumes. TEDS (tree-edit-distance similarity), introduced with PubTabNet, compares predicted and reference HTML trees, so it fits image-to-HTML and VLM table-to-markup outputs [1]. GriTS treats the table as a matrix and reports separate scores for cell topology, cell location and cell content, which helps isolate whether a failure is structural, geometric or OCR-driven. Older cell-adjacency relation metrics, used in ICDAR table competitions, check neighboring-cell pairs and are less sensitive to errors in spanning cells.

Illustrative example: invented to show structure; it does not describe an available dataset.

Downstream usePrimary metricAlso reportWhy
VLM table-to-HTML or markdown SFTTEDS (structure and content)TEDS-structure onlyMatches the serialized target the model emits
Financial spreading into a chart of accountsGriTS contentExact match on numeric cellsA wrong cell value matters more than a misplaced border
Object-detection TSR (Table Transformer style)GriTS topology and locationDetection AP on rows, columns, spansSeparates grid errors from localization errors
Regression gate for a parser releaseGriTS per hard-case sliceTEDS by document typeShows which slice regressed

Whatever you choose, score per slice (borderless, spanning headers, multi-page, scanned) and keep a held-out test set from documents never seen in training, ideally from different issuing companies.

Why born-digital documents with source files make cheaper labels

The most accurate table labels come from documents whose structured source still exists. When a report PDF was generated from a spreadsheet or a reporting system, the cell grid, spans and values can be recovered from the source file and aligned to the rendered page, the same approach PubTabNet used with XML and PDF pairs [1]. This avoids most manual cell drawing and gives exact numeric ground truth.

Ask suppliers whether they can provide the source spreadsheet, export file or report definition alongside each PDF, and whether a printed-and-scanned copy also exists. That pairing gives you a born-digital label, a scanned image of the same table, and an alignment you can verify. For native spreadsheet data as a training target in itself, see spreadsheet and financial model datasets, and for statement-specific extraction see financial statement spreading data.

A delivery schema buyers can specify

Ask for cell coordinates and logical structure delivered together with the original page image, in a documented schema. Croissant, a schema.org-based JSON-LD format, can describe the files and record structure so loaders do not depend on a README [7]. Format details for page images and labels are covered in dataset delivery formats and schemas.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "table_id": "doc_0412_p3_t1",
  "source": {"file": "doc_0412.pdf", "page": 3, "capture": "scan_300dpi", "born_digital_source": "doc_0412.xlsx"},
  "bbox": [112, 418, 1530, 1962],
  "continues_on_page": 4,
  "style": {"borders": "none", "header_levels": 2},
  "cells": [
    {"row": [0,0], "col": [1,2], "text": "Three months ended", "role": "column_header", "bbox": [640,418,1180,452]},
    {"row": [2,2], "col": [0,0], "text": "Current assets", "role": "projected_row_header", "bbox": [112,520,1530,552]},
    {"row": [3,3], "col": [1,1], "text": "(1,204)", "role": "data", "value": -1204, "bbox": [640,560,860,592]}
  ],
  "html": "<table><thead><tr><td></td><td colspan=\"2\">Three months ended</td></tr>...</thead>...</table>",
  "label_method": "aligned_from_source_spreadsheet",
  "qa": {"reviewed": true, "reviewer_agreement": null}
}

Request checklist for a TSR dataset:

  1. Document types and issuers (statements, invoices, management reports, logs) with target counts per type.
  2. Target shares for borderless, spanning-header, multi-page, scanned and photographed tables.
  3. Label layers required: detection box, grid, cells with spans, roles, text, HTML or markdown.
  4. Coordinate convention (pixel origin, DPI, PDF points) and image format.
  5. Label method per record (source-aligned, human-drawn, model-assisted then reviewed) and agreement measured on a sample.
  6. Personal data handling for names, account numbers and addresses inside cells; see de-identified data for AI training.
  7. Rights documentation for the documents themselves; see chain of title for AI training data.
  8. A held-out evaluation split by issuer, plus the metric and slices you will report.

Rights and privacy questions specific to table data

Tables in business documents concentrate sensitive values, so rights and de-identification need to be checked at the cell level. Statements and invoices carry account numbers, counterparties and amounts in cells, and replacing a value can break column totals that a model learns to check. Ask how replacements preserve formatting and arithmetic, and which fields were left intact.

Public datasets carry their own constraints: confirm the license of each set you train on, since several were built from third-party publications and release terms vary. Licensed real documents should arrive with a license that states the records covered, permitted uses and delivery terms.

How SourceX approaches table extraction data

SourceX sources operational datasets from US companies on request, including documents and finance and legal workflow records, and manages licensing and ongoing purchases. Nothing is held in stock, and a request does not guarantee a match. Buyers describe the tables and hard cases they need, SourceX looks for US businesses that hold that data, and every release is approved by the supplying company. You can start by describing your table extraction data requirements.

Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, account numbers and emails are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. For the wider task map, see document AI datasets by task and the AI data hub; for RAG use of the same documents, see RAG evaluation datasets from real company documents.

Request real-document table structure recognition data

Describe the document types, table styles and label layers you need, and SourceX will look for US companies that hold matching data and assess their licensing permissions. Pricing and allowed uses are agreed in a license, and nothing is contracted until a supplier agrees. Describe the table data you need.

Sources

  1. arXiv (Zhong, ShafieiBavani, Jimeno Yepes), "Image-based table recognition: data, model, and evaluation (PubTabNet)" (2019). https://arxiv.org/pdf/1911.10683v5
  2. arXiv (Li et al.), "TableBank: A Benchmark Dataset for Table Detection and Recognition" (2019). https://arxiv.org/pdf/1903.01949
  3. CVF Open Access (WACV 2025), "CISOL: An Open and Extensible Dataset for Table Structure Recognition in the Construction Industry" (2025). https://openaccess.thecvf.com/content/WACV2025/html/Tschirschwitz_CISOL_An_Open_and_Extensible_Dataset_for_Table_Structure_Recognition_WACV_2025_paper.html
  4. arXiv, "SynFinTabs: A Dataset of Synthetic Financial Tables for Information and Table Extraction" (2024). https://arxiv.org/pdf/2412.04262
  5. arXiv, "TableNet: A Large-Scale Table Dataset with LLM-Powered Autonomous generation" (2026). https://arxiv.org/pdf/2604.13041
  6. arXiv (Jaume, Ekenel, Thiran), "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
  7. arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data