Skip to content

Document AI data

Document Layout Analysis Datasets: Going Beyond DocLayNet and PubLayNet

Quick answer

The standard document layout analysis datasets are PubLayNet, DocBank and DocLayNet, with FUNSD covering scanned forms. PubLayNet and DocBank were labeled automatically from scientific publications, and DocLayNet added human annotation across more document types, but none of the three contains the invoices, bills of lading, stamped correspondence and fax-degraded forms that dominate enterprise intake. Use the public sets for pretraining and benchmarking, then source or commission layout-annotated pages whose document mix and class schema match production.

By SourceX Editorial · Updated

What the public layout benchmarks actually contain

Each public benchmark reflects how it was built, and that construction method sets its ceiling for business documents. PubLayNet was generated by matching XML with the content of over 1 million PubMed Central PDF articles, producing over 360 thousand page images, and its authors report that detectors trained on it recognize scientific article layouts accurately [1]. DocBank took a similar route with 500K pages labeled at token level by weak supervision from arXiv LaTeX sources [2]. Both are born-digital, single-domain and clean.

DocLayNet was built because those sets lacked layout variability. It contains 80,863 manually annotated pages with bounding boxes in 11 classes, distributed in COCO format, and a subset was annotated two or three times to measure inter-annotator agreement [3]. The paper reports that baseline detectors scored roughly 10% below that agreement, which tells you the human ceiling and the realistic headroom [3]. A later DocLayNet release is hosted on Hugging Face, so record which version your numbers come from.

FUNSD is the form-oriented outlier: 199 noisy, low-resolution scanned forms annotated for text detection, OCR, spatial layout, and entity labeling and linking [4]. It is valuable as a scanned-form test, but at 199 pages it is an evaluation set, not a training corpus.

DatasetSource documentsLabeling methodSizeFit for enterprise layout
PubLayNet [1]PubMed Central articlesAutomatic, XML-to-PDF matching360K+ pagesPretraining only; scientific layouts
DocBank [2]arXiv papersWeak supervision from LaTeX500K pagesToken-level pretraining; academic domain
DocLayNet [3]Diverse born-digital sourcesHuman, with agreement subset80,863 pagesBest public baseline; no forms, stamps or handwriting classes
FUNSD [4]Scanned formsHuman199 formsNoisy-scan evaluation; too small to train
Ultralytics signature [6]Document imagesSingle signature class178 imagesSignature-region smoke test only

Where publication-derived layouts fail on business documents

Detectors trained on publication-style data miss the regions that carry business meaning. An article page is a predictable grid of title, paragraphs, figures, captions and tables. A commercial invoice has a logo block, a remit-to address, a key-value header, a line-item table that spans pages, a totals box, a payment stub and sometimes a handwritten approval.

The recurring failure modes are concrete:

  • Missing classes. Stamps, signatures, handwritten annotations, barcodes, checkboxes and key-value blocks have no class in a five- or eleven-class publication schema, so the detector folds them into Text or Figure/Picture.
  • Capture gap. Born-digital PDFs rendered at a fixed resolution do not teach the model skew, fax banding, punch-hole shadows or phone-photo perspective (see degraded document capture data).
  • Density gap. Forms pack dozens of small fields per page; detectors tuned on paragraph-sized regions under-detect small boxes.
  • Template leakage. Business documents repeat vendor templates, so random page splits put near-identical layouts in both train and test and inflate mAP. Near-duplicate inflation is well documented for text corpora [9]; split by issuer or template ID instead.
  • Multi-page context. Continuation tables and repeated headers only make sense across pages, which single-page benchmarks never test (see multi-page document data).

A single-class set such as the Ultralytics signature data, with 178 images split 143/35 [6], shows how thin the public supply is for business-specific regions.

Designing the class schema before you source pages

Fix the class list, geometry and nesting rules before collecting a page, because they determine which documents are worth acquiring. A practical approach keeps DocLayNet's 11 classes as a base so you can still benchmark against it [3], then adds business classes as children or siblings. Our document annotation schema guide covers reading order, tables and fields in one spec; this page focuses on the layout layer.

Decisions to write down:

  • Geometry. Axis-aligned boxes are cheapest and suit most detectors; polygons matter for skewed scans, stamps over text and rotated photos. COCO supports both, plus RLE masks [5].
  • Nesting. Decide whether a key-value block contains its key and value regions, and whether a table region coexists with cell-level structure from a table structure recognition set.
  • Overlap policy. Stamps and signatures often overlap printed text; specify whether both regions are labeled.
  • Agreement. Double-annotate a fixed share of pages and report a metric such as box-level F1 at IoU 0.5 per class, following DocLayNet's practice of measuring human agreement before training [3].

Illustrative example: invented to show structure; it does not describe an available dataset.

ClassParentGeometryRule
Page-header / Page-footerpageboxRepeated running content only
Section-header, Title, Text, List-itempageboxDocLayNet-compatible
TablepageboxLine items including column headers; cells in separate layer
Key-value-blockpageboxLabel and value pairs grouped as one field cluster
SignaturepagepolygonInk only, excluding printed name line
StamppagepolygonMay overlap Text; both labeled
Handwritten-notepagepolygonMarginal or inline handwriting not part of a field
Barcode / CheckboxpageboxCheckbox carries checked attribute

Formats and metadata that keep layout data usable

COCO-style JSON is the default exchange format for layout detectors, and the details that matter are resolution and provenance. A COCO file holds info, licenses, categories, images and annotations sections, with boxes as [x, y, width, height] in absolute pixels and segmentations as polygons or RLE [5]. Because coordinates are pixel-based, keep page images at their original DPI and never resample after labeling.

Add per-image fields that COCO does not define: whether the page is scanned or born-digital, capture DPI, page index within the source document, document type and an issuer or template ID for grouped splits. Document the dataset with a Data Card covering sources, annotation method and known gaps [7], and ship machine-readable Croissant metadata so loaders understand the file and record structure [8].

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "images": [{
    "id": 4107, "file_name": "inv_000412_p1.png", "width": 2550, "height": 3300,
    "dpi": 300, "capture": "scanned", "doc_type": "commercial_invoice",
    "doc_id": "D-000412", "page_index": 1, "page_count": 2, "template_group": "T-0193"
  }],
  "categories": [{"id": 12, "name": "Key-value-block"}, {"id": 14, "name": "Stamp"}],
  "annotations": [
    {"id": 1, "image_id": 4107, "category_id": 12, "bbox": [180, 410, 960, 380], "iscrowd": 0},
    {"id": 2, "image_id": 4107, "category_id": 14,
     "segmentation": [[1820, 2900, 2210, 2880, 2230, 3120, 1840, 3140]], "iscrowd": 0}
  ]
}

Mixing public sets with licensed business pages

Public layout sets carry their own license terms, so check commercial-use permissions for each one before blending it into a training mix. Several research datasets restrict commercial use or inherit terms from the underlying documents, and a mixed corpus is only as usable as its most restrictive component; our guide to public document datasets and commercial use walks through the common ones.

A common recipe is to pretrain on PubLayNet or DocBank for generic text-block detection, fine-tune on DocLayNet for class diversity, then fine-tune and evaluate on licensed pages from your own target domains. Keep the evaluation split strictly from the business pages, grouped by issuer, so reported gains reflect production. If you are weighing generated pages, read synthetic vs real documents first; synthetic forms rarely reproduce real stamps, overprinting and scan artifacts.

Licensed business pages raise privacy questions that public benchmarks avoid. Invoices, statements and correspondence carry names, addresses and account numbers, and layout labels do not remove them; plan redaction or replacement before delivery and confirm that redaction boxes are not themselves learned as a layout class (see de-identified data for AI training).

A request template for layout-annotated business pages

The most reliable way to get relevant pages is to describe the documents and the labels, not the companies that might hold them. A request that names document types, capture conditions, class schema and acceptance tests lets a supplier decide quickly whether its archive fits.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample entry
Document types and mixCommercial invoices 40%, purchase orders 20%, packing lists 15%, signed delivery receipts 15%, remittance advices 10%
CaptureAt least 50% scanned at 200-300 DPI; include fax and phone-photo pages; record capture per page
Language and scriptUS English; flag mixed-script pages
Class schemaDocLayNet 11 classes plus Key-value-block, Signature, Stamp, Handwritten-note, Barcode, Checkbox
GeometryBoxes, polygons for Signature, Stamp, Handwritten-note
FormatCOCO JSON per split, original-resolution PNG or TIFF, per-image provenance fields
Quality10% double-annotated; report per-class box F1 at IoU 0.5
SplitsGrouped by issuer and template; no template across splits
PrivacyPersonal details removed or replaced; method documented
Intended useTrain and evaluate a layout detector for a parsing pipeline

The same pages often support adjacent tasks, such as PDF parsing to Markdown or HTML, key-value extraction labels and document classification, so ask whether text, reading order or field labels can be added under the same license. The Document AI data hub maps those annotation layers together.

How SourceX approaches layout data requests

SourceX sources operational datasets, including business documents, from US companies on request and manages the licensing process; it does not hold layout datasets in stock, and a request does not guarantee a match. Buyers describe the data they need, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery.

Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. SourceX does not source scraped web content and does not train models. For broader context, see training data for document understanding models and enterprise document datasets for AI training, or start a request on the SourceX buyer page.

Sourcing layout-annotated business documents

If public benchmarks leave gaps in your document mix or class schema, describe the pages, capture conditions and labels you need. SourceX looks for US companies that hold matching documents, assesses data and licensing permissions, and nothing is contracted until a supplier agrees. Describe the layout data you need.

Sources

  1. arXiv (Zhong, Tang, Jimeno Yepes), "PubLayNet: largest dataset ever for document layout analysis" (2019). https://arxiv.org/pdf/1908.07836
  2. COLING 2020 (ACL Anthology), "DocBank: A Benchmark Dataset for Document Layout Analysis" (2020). https://preview.aclanthology.org/setup/2020.coling-main.82
  3. arXiv / KDD 2022 (Pfitzmann et al., IBM Research), "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
  4. arXiv (Jaume, Ekenel, Thiran), "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
  5. CVAT.ai, "COCO format documentation". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
  6. Ultralytics, "Signature Detection Dataset". https://docs.ultralytics.com/datasets/detect/signature
  7. FAccT 2022 / arXiv (Pushkarna, Zaldivar, Kjartansson), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  8. arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  9. ACL 2022 / arXiv (Lee et al.), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data