Document AI data
Document Layout Analysis Datasets: Going Beyond DocLayNet and PubLayNet
Quick answer
The standard document layout analysis datasets are PubLayNet, DocBank and DocLayNet, with FUNSD covering scanned forms. PubLayNet and DocBank were labeled automatically from scientific publications, and DocLayNet added human annotation across more document types, but none of the three contains the invoices, bills of lading, stamped correspondence and fax-degraded forms that dominate enterprise intake. Use the public sets for pretraining and benchmarking, then source or commission layout-annotated pages whose document mix and class schema match production.
By SourceX Editorial · Updated
What the public layout benchmarks actually contain
Each public benchmark reflects how it was built, and that construction method sets its ceiling for business documents. PubLayNet was generated by matching XML with the content of over 1 million PubMed Central PDF articles, producing over 360 thousand page images, and its authors report that detectors trained on it recognize scientific article layouts accurately [1]. DocBank took a similar route with 500K pages labeled at token level by weak supervision from arXiv LaTeX sources [2]. Both are born-digital, single-domain and clean.
DocLayNet was built because those sets lacked layout variability. It contains 80,863 manually annotated pages with bounding boxes in 11 classes, distributed in COCO format, and a subset was annotated two or three times to measure inter-annotator agreement [3]. The paper reports that baseline detectors scored roughly 10% below that agreement, which tells you the human ceiling and the realistic headroom [3]. A later DocLayNet release is hosted on Hugging Face, so record which version your numbers come from.
FUNSD is the form-oriented outlier: 199 noisy, low-resolution scanned forms annotated for text detection, OCR, spatial layout, and entity labeling and linking [4]. It is valuable as a scanned-form test, but at 199 pages it is an evaluation set, not a training corpus.
| Dataset | Source documents | Labeling method | Size | Fit for enterprise layout |
|---|---|---|---|---|
| PubLayNet [1] | PubMed Central articles | Automatic, XML-to-PDF matching | 360K+ pages | Pretraining only; scientific layouts |
| DocBank [2] | arXiv papers | Weak supervision from LaTeX | 500K pages | Token-level pretraining; academic domain |
| DocLayNet [3] | Diverse born-digital sources | Human, with agreement subset | 80,863 pages | Best public baseline; no forms, stamps or handwriting classes |
| FUNSD [4] | Scanned forms | Human | 199 forms | Noisy-scan evaluation; too small to train |
| Ultralytics signature [6] | Document images | Single signature class | 178 images | Signature-region smoke test only |
Where publication-derived layouts fail on business documents
Detectors trained on publication-style data miss the regions that carry business meaning. An article page is a predictable grid of title, paragraphs, figures, captions and tables. A commercial invoice has a logo block, a remit-to address, a key-value header, a line-item table that spans pages, a totals box, a payment stub and sometimes a handwritten approval.
The recurring failure modes are concrete:
- Missing classes. Stamps, signatures, handwritten annotations, barcodes, checkboxes and key-value blocks have no class in a five- or eleven-class publication schema, so the detector folds them into Text or Figure/Picture.
- Capture gap. Born-digital PDFs rendered at a fixed resolution do not teach the model skew, fax banding, punch-hole shadows or phone-photo perspective (see degraded document capture data).
- Density gap. Forms pack dozens of small fields per page; detectors tuned on paragraph-sized regions under-detect small boxes.
- Template leakage. Business documents repeat vendor templates, so random page splits put near-identical layouts in both train and test and inflate mAP. Near-duplicate inflation is well documented for text corpora [9]; split by issuer or template ID instead.
- Multi-page context. Continuation tables and repeated headers only make sense across pages, which single-page benchmarks never test (see multi-page document data).
A single-class set such as the Ultralytics signature data, with 178 images split 143/35 [6], shows how thin the public supply is for business-specific regions.
Designing the class schema before you source pages
Fix the class list, geometry and nesting rules before collecting a page, because they determine which documents are worth acquiring. A practical approach keeps DocLayNet's 11 classes as a base so you can still benchmark against it [3], then adds business classes as children or siblings. Our document annotation schema guide covers reading order, tables and fields in one spec; this page focuses on the layout layer.
Decisions to write down:
- Geometry. Axis-aligned boxes are cheapest and suit most detectors; polygons matter for skewed scans, stamps over text and rotated photos. COCO supports both, plus RLE masks [5].
- Nesting. Decide whether a key-value block contains its key and value regions, and whether a table region coexists with cell-level structure from a table structure recognition set.
- Overlap policy. Stamps and signatures often overlap printed text; specify whether both regions are labeled.
- Agreement. Double-annotate a fixed share of pages and report a metric such as box-level F1 at IoU 0.5 per class, following DocLayNet's practice of measuring human agreement before training [3].
Illustrative example: invented to show structure; it does not describe an available dataset.
| Class | Parent | Geometry | Rule |
|---|---|---|---|
| Page-header / Page-footer | page | box | Repeated running content only |
| Section-header, Title, Text, List-item | page | box | DocLayNet-compatible |
| Table | page | box | Line items including column headers; cells in separate layer |
| Key-value-block | page | box | Label and value pairs grouped as one field cluster |
| Signature | page | polygon | Ink only, excluding printed name line |
| Stamp | page | polygon | May overlap Text; both labeled |
| Handwritten-note | page | polygon | Marginal or inline handwriting not part of a field |
| Barcode / Checkbox | page | box | Checkbox carries checked attribute |
Formats and metadata that keep layout data usable
COCO-style JSON is the default exchange format for layout detectors, and the details that matter are resolution and provenance. A COCO file holds info, licenses, categories, images and annotations sections, with boxes as [x, y, width, height] in absolute pixels and segmentations as polygons or RLE [5]. Because coordinates are pixel-based, keep page images at their original DPI and never resample after labeling.
Add per-image fields that COCO does not define: whether the page is scanned or born-digital, capture DPI, page index within the source document, document type and an issuer or template ID for grouped splits. Document the dataset with a Data Card covering sources, annotation method and known gaps [7], and ship machine-readable Croissant metadata so loaders understand the file and record structure [8].
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"images": [{
"id": 4107, "file_name": "inv_000412_p1.png", "width": 2550, "height": 3300,
"dpi": 300, "capture": "scanned", "doc_type": "commercial_invoice",
"doc_id": "D-000412", "page_index": 1, "page_count": 2, "template_group": "T-0193"
}],
"categories": [{"id": 12, "name": "Key-value-block"}, {"id": 14, "name": "Stamp"}],
"annotations": [
{"id": 1, "image_id": 4107, "category_id": 12, "bbox": [180, 410, 960, 380], "iscrowd": 0},
{"id": 2, "image_id": 4107, "category_id": 14,
"segmentation": [[1820, 2900, 2210, 2880, 2230, 3120, 1840, 3140]], "iscrowd": 0}
]
}
Mixing public sets with licensed business pages
Public layout sets carry their own license terms, so check commercial-use permissions for each one before blending it into a training mix. Several research datasets restrict commercial use or inherit terms from the underlying documents, and a mixed corpus is only as usable as its most restrictive component; our guide to public document datasets and commercial use walks through the common ones.
A common recipe is to pretrain on PubLayNet or DocBank for generic text-block detection, fine-tune on DocLayNet for class diversity, then fine-tune and evaluate on licensed pages from your own target domains. Keep the evaluation split strictly from the business pages, grouped by issuer, so reported gains reflect production. If you are weighing generated pages, read synthetic vs real documents first; synthetic forms rarely reproduce real stamps, overprinting and scan artifacts.
Licensed business pages raise privacy questions that public benchmarks avoid. Invoices, statements and correspondence carry names, addresses and account numbers, and layout labels do not remove them; plan redaction or replacement before delivery and confirm that redaction boxes are not themselves learned as a layout class (see de-identified data for AI training).
A request template for layout-annotated business pages
The most reliable way to get relevant pages is to describe the documents and the labels, not the companies that might hold them. A request that names document types, capture conditions, class schema and acceptance tests lets a supplier decide quickly whether its archive fits.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example entry |
|---|---|
| Document types and mix | Commercial invoices 40%, purchase orders 20%, packing lists 15%, signed delivery receipts 15%, remittance advices 10% |
| Capture | At least 50% scanned at 200-300 DPI; include fax and phone-photo pages; record capture per page |
| Language and script | US English; flag mixed-script pages |
| Class schema | DocLayNet 11 classes plus Key-value-block, Signature, Stamp, Handwritten-note, Barcode, Checkbox |
| Geometry | Boxes, polygons for Signature, Stamp, Handwritten-note |
| Format | COCO JSON per split, original-resolution PNG or TIFF, per-image provenance fields |
| Quality | 10% double-annotated; report per-class box F1 at IoU 0.5 |
| Splits | Grouped by issuer and template; no template across splits |
| Privacy | Personal details removed or replaced; method documented |
| Intended use | Train and evaluate a layout detector for a parsing pipeline |
The same pages often support adjacent tasks, such as PDF parsing to Markdown or HTML, key-value extraction labels and document classification, so ask whether text, reading order or field labels can be added under the same license. The Document AI data hub maps those annotation layers together.
How SourceX approaches layout data requests
SourceX sources operational datasets, including business documents, from US companies on request and manages the licensing process; it does not hold layout datasets in stock, and a request does not guarantee a match. Buyers describe the data they need, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery.
Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. SourceX does not source scraped web content and does not train models. For broader context, see training data for document understanding models and enterprise document datasets for AI training, or start a request on the SourceX buyer page.
Sourcing layout-annotated business documents
If public benchmarks leave gaps in your document mix or class schema, describe the pages, capture conditions and labels you need. SourceX looks for US companies that hold matching documents, assesses data and licensing permissions, and nothing is contracted until a supplier agrees. Describe the layout data you need.
Sources
- arXiv (Zhong, Tang, Jimeno Yepes), "PubLayNet: largest dataset ever for document layout analysis" (2019). https://arxiv.org/pdf/1908.07836
- COLING 2020 (ACL Anthology), "DocBank: A Benchmark Dataset for Document Layout Analysis" (2020). https://preview.aclanthology.org/setup/2020.coling-main.82
- arXiv / KDD 2022 (Pfitzmann et al., IBM Research), "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
- arXiv (Jaume, Ekenel, Thiran), "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
- CVAT.ai, "COCO format documentation". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
- Ultralytics, "Signature Detection Dataset". https://docs.ultralytics.com/datasets/detect/signature
- FAccT 2022 / arXiv (Pushkarna, Zaldivar, Kjartansson), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- ACL 2022 / arXiv (Lee et al.), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.