Document AI data
Document AI Datasets by Task: Annotation Layers, Public Sets and Licensed Business Documents
Quick answer
Document AI datasets pair page images or PDFs with annotation layers: OCR transcriptions, layout regions, reading order, table structure, key-value fields, document classes or question-answer pairs. Choose the layer by task, then the source: public research sets such as DocLayNet and FUNSD for benchmarking and pre-training where their terms allow, synthetic pages for volume, and licensed business documents when a model must handle real issuers, poor scans, handwriting and multi-page packets.
By SourceX Editorial · Updated
Nine document AI tasks and the annotation layer each one needs
Every document AI task learns from a different annotation layer, so fix the layer, labeling unit and metric before comparing datasets or suppliers. Optical character recognition (OCR) sits underneath most rows.
| Task | Annotation layer to specify | Common container | Score with | Go deeper |
|---|---|---|---|---|
| Printed OCR and handwriting recognition (HTR) | Text aligned to word or line boxes, with handwriting and legibility flags | ALTO XML, hOCR or PAGE XML; JSON tokens | Character and word error rate (CER, WER) | OCR ground truth; handwriting data |
| Layout analysis | Region boxes or polygons with classes (title, text, table, picture, footnote) | COCO-style JSON | Mean average precision across overlap thresholds | Layout datasets beyond DocLayNet |
| Reading order | Ordered regions or lines across columns, sidebars and form flows | Order index per region | Edit distance to the gold sequence | Reading order data |
| Table structure recognition | Cell grid with row and column spans, header flags and cell text | HTML or JSON cell lists | Tree-edit-distance similarity (TEDS) or cell-adjacency F1 | Table structure data |
| Key-value and line-item extraction | Field name, value linked to source tokens, normalized value, line-item grouping | JSON with token references | Field-level precision and recall after normalization | Key-value labels; invoice line items |
| Classification and splitting | Class per document; boundary labels inside a scanned packet | Per-page CSV or JSON | Per-class F1; boundary F1 | Classification data; page-stream segmentation |
| Document QA and VLM fine-tuning | Question, answer, evidence page and region; page-to-Markdown or HTML targets | JSONL | Exact match or normalized string similarity; evidence accuracy | Document QA data; PDF parsing data |
| Charts and drawings | Chart paired with its source table; title-block fields, dimensions, symbols | Images with JSON or CSV | Value-level accuracy; symbol detection precision and recall | Chart data; engineering drawings |
| Redaction, signatures and stamps | Sensitive spans and boxes with human redaction decisions; signature, stamp and seal boxes | Span lists plus COCO-style boxes | Recall per entity type; detection precision | Redaction data; signature and stamp data |
Layers stack: one OCR token box is the evidence for an extracted field, a table cell's content and a question's answer region, so buy layers in one schema keyed to one coordinate frame (one annotation spec for layout, reading order, tables and fields). For layout, DocLayNet ships 80,863 manually annotated pages with boxes in 11 classes in COCO format [1]. The COCO annotation format's JSON specification includes fields for license information [2].
One page, every layer: what a deliverable record contains
A deliverable record keeps the page image, every annotation layer, each label's origin, the license reference and the privacy treatment together, so a reviewer can trace any value back to pixels and a license.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"doc_id": "pkt-20931-d02",
"packet": {"packet_id": "pkt-20931", "pages": [3, 4], "doc_class": "vendor_invoice",
"packet_classes": ["cover_letter", "vendor_invoice", "proof_of_delivery"]},
"page": {"page_no": 3, "image": "pages/pkt-20931-p003.png", "dpi": 300,
"width_px": 2550, "height_px": 3300, "rotation_deg": 0,
"capture": "fax_bw", "box_convention": "x0_y0_x1_y1_px_top_left"},
"ocr_tokens": [
{"id": "t118", "text": "INV-55210", "bbox": [1890, 212, 2210, 258], "conf": 0.97, "style": "printed"},
{"id": "t119", "text": "Net 30", "bbox": [1890, 300, 2040, 344], "conf": 0.71, "style": "handwritten"}
],
"layout_regions": [
{"id": "r2", "class": "key_value_block", "bbox": [1700, 180, 2400, 360], "reading_order": 2},
{"id": "r6", "class": "table", "bbox": [150, 1120, 2400, 2280], "reading_order": 6}
],
"table_cells": [
{"region": "r6", "row": 2, "col": 3, "row_span": 1, "col_span": 1, "is_header": false, "tokens": ["t402"]}
],
"fields": [
{"key": "invoice_number", "value": "INV-55210", "tokens": ["t118"], "source": "erp_ap_entry"},
{"key": "payment_terms", "value": "Net 30", "normalized": "NET_30", "tokens": ["t119"],
"source": "annotator_double_pass"}
],
"rights": {"license_ref": "LIC-0142", "third_parties": ["issuing_vendor"],
"permitted_uses": ["model_training", "internal_evaluation"]},
"privacy": {"actions": ["signature_masked", "remit_account_replaced_with_surrogate"],
"layers_cleaned": ["pixels", "text_layer", "pdf_metadata"], "method_ref": "DEID-V3"},
"split_key": "issuer:9f2c"
}
Check four things in a record like this: a declared box convention (units, corner or width-height form, origin), since mixed conventions silently misalign labels; values linked to OCR token IDs, so grounding can be scored; a source field separating system-of-record values, such as an accounts-payable entry, from annotator labels; and a split key that keeps each issuer or template on one side of a split.
Public sets, generated pages and licensed archives compared
Most document AI programs combine three routes, each failing a different check: public datasets on coverage and license clarity, synthetic pages on realism, and licensed business documents on per-deal rights and privacy work.
| Route | Strong for | Weak for | Check before use |
|---|---|---|---|
| Public research datasets | Layout pre-training, benchmarks, baselines | Business forms at volume, scan noise, line items, clear commercial terms | License scope (pages or annotations); release; test-set overlap |
| Synthetic or rendered documents | Volume, rare field combinations, no personal data | Issuer and template diversity, real degradation, handwriting, stamps, packet structure | Accuracy on issuers the generator never saw |
| Licensed business documents | Real issuers, template versions, capture conditions, operational labels | Speed; per-deal rights review and de-identification | Holder's right to share; third-party content; redaction method |
Public research datasets. Models trained on DocLayNet were more robust than those trained on PubLayNet or DocBank, which come mainly from scientific repositories [1]; a later DocLayNet v1.2 is hosted on Hugging Face by the docling-project [3], so name the release and read its card. FUNSD's 199 noisy scanned forms, annotated for text detection, OCR, layout and entity labeling and linking [4], can benchmark form understanding, but the paper notes the dataset may not be large enough to create a generalizable application. EDGAR-CORPUS, US 10-K reports from 1993 to 2020 split into items in JSON [5], helps language pre-training but teaches nothing about layout.
License metadata is unreliable. The Data Provenance Initiative's audit of more than 1,800 text datasets found license omissions above 70% and error rates above 50% on popular hosting sites [6], and a 2021 study found potential license-violation risks in five of six widely used public image datasets used commercially, partly because one dataset can mix sources under different licenses [7]. Check whether a license covers the page images or only the annotations (which public document datasets allow commercial use).
Synthetic documents. Template rendering gives volume and free labels, but a model can learn where a template puts a field instead of what the field means. Test on issuers and capture conditions the generator never produced, such as faxes, photocopies and phone photos (real-world capture conditions; where synthetic documents break).
Licensed business documents. Operational archives hold what generators rarely reproduce: many issuers, template revisions over years, margin notes, stamps over print and packets that mix document types. They often hold labels too, such as values keyed into an ERP (documents paired with system-of-record entries) and reviewers' corrections to production extraction (IDP correction logs). Compare annotating your own pages with licensing labeled ones.
SourceX sources operational datasets from US companies, including documents and finance and legal workflow records, and manages the licensing agreement. Buyers describe the documents; SourceX looks for businesses that hold them, and the supplying company approves every release. These are kinds of data it sources, not inventory under contract, so a request does not guarantee a match, and it does not source scraped public web content. See enterprise document datasets, training data for multimodal document models and the document AI and RAG use case, or describe your document requirements to SourceX.
Who else is on the page: counterparties, patients and account holders
A business document rarely belongs to one party: the company that holds it, the counterparty that issued it, and the people named in it all have interests. Check rights and privacy per document type, not per archive.
- Counterparty terms. Contracts, supplier invoices and bills of lading name another company and its prices and terms; ask whether the holder's confidentiality clauses restrict sharing and which values to mask.
- Health documents. CMS-1500 and UB-04 claim forms, explanations of benefits and prior-authorization faxes carry protected health information. HIPAA de-identification uses Safe Harbor, removing 18 listed identifiers of the individual and of relatives, employers or household members with no actual knowledge of identifiability, or an expert determination that identification risk is very small [8]. For health records, SourceX requires one of these methods before anything is considered for a license.
- Financial documents. Bank statements, pay stubs and loan files held by lenders carry nonpublic personal information. Under Regulation P (12 CFR 1016.11), a recipient that gets it from a nonaffiliated financial institution under an exception may use and disclose it only in the ordinary course of business to carry out the purpose for which it was received [9], so ask how the holder obtained it and what its own privacy notice allows it to share (bank statements and income documents).
- Pixels, text layers and metadata. A PDF holds personal data in rendered pixels, the text or OCR layer and metadata; masking one leaves the others readable. Presidio, an open-source PII detection and anonymization SDK, states there is no guarantee it finds all sensitive information [10] (redacting PII in scanned documents). Black boxes also teach a model black boxes, so ask which values were replaced in place with realistic surrogates.
- Published PDFs for pre-training. For general-purpose models placed on the EU market, Article 53 of the AI Act requires a copyright policy that honors text-and-data-mining reservations and a public summary of training content [11] (licensed PDF corpora for pre-training).
On deals SourceX manages, every dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place; the license defines included records, allowed uses, term and delivery. Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded per dataset and a sample is checked after processing; no de-identification method is perfect. Compare methods in the de-identified data buyer's guide.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Scoring document models on documents they have never seen
Evaluate document models on a private held-out set split by issuer, template and time and scored at field level, because public benchmarks may already sit in pre-training data and repeated templates inflate scores.
Contamination is the first reason. A 2024 study notes that many large language models' training data is contaminated with test data, and that private holdout sets that could verify scores do not exist for most benchmarks [12]. Assume a document VLM may have seen public document benchmarks.
Template repetition is the second: one vendor can send thousands of invoices from one template. Language-model corpora show the same pattern: they contain many near-duplicate examples, and models trained on deduplicated versions emitted memorized text about ten times less often [13]. Deduplicate near-identical pages, then keep each issuer and template family on one side of the split.
Gold labels need an audit too. Test sets of 10 widely used vision, language and audio datasets had an estimated average label error rate of at least 3.3%, and such errors can change which model ranks best [14]. Double-annotate a subset, as DocLayNet did [1], and report Krippendorff's alpha, where 1 is perfect reliability and 0 is none beyond chance [15].
Write scoring rules into the specification: date, currency and amount normalization; scoring for fields absent from the page; line-item matching when rows split or merge; and whether a correct value with the wrong evidence box counts. See document extraction evaluation sets, private evaluation sets vs public benchmarks and evaluation datasets from real business work.
Start here: business document types and their guides
Each record type links to its cluster guides and to the SourceX page on licensing it.
| Record type | Cluster guides | SourceX page |
|---|---|---|
| Invoices, purchase orders, receipts | Invoice line items; receipts; three-way match | Invoices and receipts; purchase orders |
| Scanned and handwritten forms | ACORD forms; checkbox and selection marks | Scanned forms and handwriting; underwriting files |
| Financial and bank statements | Financial statement spreading; bank statements | Finance and accounting AI data |
| Contracts and agreements | Contract clause annotation; contract families | Contract redlines |
| Insurance and health claims | Loss run extraction | Insurance claims; healthcare revenue cycle |
| Freight and trade documents | Freight documents; trade documents; trade finance | Supply chain and logistics |
| Drawings and diagrams | Floor plans; P&IDs | CAD and PCB design files |
| Quality and safety documents | Certificates of analysis; safety data sheets | Manufacturing quality records |
| Mixed archives and long PDFs | Licensed PDF corpora; long documents; multilingual documents | Enterprise document archives |
Also see document fraud detection data, document version pairs, the multimodal training data hub and the AI data buyer's guide.
Eight lines to put in a document data request
A document data request is complete when a supplier can tell from it which pages qualify, which layers to deliver and how you will accept them.
- Tasks and layers. The tasks from the first table and their layers, in one schema.
- Document mix. Record types, issuers or templates per type, pages per document, and whether packets stay intact.
- Capture conditions. Born-digital PDFs versus scans, faxes and phone photos; resolution; share of pages with handwriting, stamps or signatures.
- Languages and scripts. Including mixed-script and right-to-left pages.
- Label origin and quality. System-of-record values or annotator passes, agreement targets and adjudication.
- Splits and holdouts. Issuer- or template-disjoint splits and an evaluation slice that never enters training.
- Privacy and rights. Fields masked or replaced across pixels, text layers and metadata, and the uses you need licensed.
- Documentation and delivery. A datasheet covering motivation, composition, collection process and recommended uses [16], plus formats and transfer (delivery formats and transfer).
The document dataset requirements spec turns this list into a full request.
Looking for real business documents for document AI?
Describe the document types, annotation layers, volumes, capture conditions and allowed uses you need on the SourceX buyer page. SourceX looks for US companies that hold those documents, checks the data and the supplier's licensing permissions, manages the license and coordinates delivery and payment; nothing is contracted until a supplier agrees. Describe the documents your model needs.
Guides in this section
- ACORD Form Extraction Data for Submission Intake AIHow to source filled ACORD 125, 126 and 140 forms and full submission packets with field labels for insurance intake extraction and triage models.
- Bank Statement and Pay Stub Datasets for Lending Document AISourcing bank statements, pay stubs and W-2s for lending document AI: transaction labels, format coverage, GLBA and tax-preparer limits, synthetic gaps
- Contract Clause Extraction Datasets: Labels, Spans, RightsHow to source contract clause extraction data: CUAD limits, clause taxonomy mapping, absent-clause rules, span vs presence metrics and licensing checks.
- Document AI Datasets: Which Allow Commercial Use?License and provenance status of FUNSD, RVL-CDIP, IIT-CDIP, DocVQA, DocLayNet and PubLayNet for commercial document AI training, as of October 2026.
- Document Classification Datasets from Real Intake MixesSource document classification training data that matches your intake: taxonomy, natural class mix, other/unknown class, page vs document labels.
- Financial Statement Extraction Data for Spreading ModelsHow to source financial statement extraction data: real PDFs paired with analyst spreads, coverage to request, XBRL as a complement, and validation checks.
- Handwriting Recognition Training Data for Business FormsHow to source handwriting recognition training data from real forms and notes: writer diversity, field crops, transcription rules, writer-disjoint tests.
- Invoice Line-Item Extraction Data: Vendors and TemplatesHow to source invoice line-item extraction data: vendor and template diversity, multi-page line tables, posted AP labels and held-out evaluation splits.
- Key Information Extraction Datasets: Specifying KV LabelsHow to specify key information extraction data: field values, boxes, key-value links and normalized values on real business forms, beyond FUNSD and DocILE.
- Licensed PDF Corpora for VLM and Multimodal Pre-TrainingHow AI labs license non-public business PDFs for VLM pre-training: composition reports, scanned vs born-digital mix, deduplication and rights terms.
- OCR Ground Truth Data for Business Scans: Buyer's GuideHow to specify and license OCR ground truth for business scans: word and line alignment, ALTO, PAGE XML and hOCR formats, CER acceptance and rights.
- PDF to Markdown Training Data: Page Images and Markup LabelsHow to source PDF parsing training data: page images paired with Markdown or HTML labels, native-source pairs, markup conventions and evaluation metrics.
- Receipt Line-Item Data for Expense AI Training and EvalHow to source receipt photos with line-item, tax and total labels: limits of CORD and SROIE, capture conditions, expense-report labels and redaction.
- Synthetic vs Real Documents for Document AI TrainingWhere synthetic document data works for OCR and layout pre-training, where it breaks on template diversity, handwriting and degradation, and how to mix it.
- Table Structure Recognition Data from Real DocumentsWhat table structure recognition datasets cover, where PubTabNet, PubTables-1M and FinTabNet fall short, and how to specify real business table data.
- Chart Understanding Data: Real Charts with Source TablesHow to source chart understanding data: real business charts paired with their exact source tables and QA, beyond synthetic and web-sourced public sets.
- Commercial Invoice and Customs Document Extraction DataSourcing commercial invoices, packing lists and certificates of origin with field labels for customs entry and trade-compliance document AI.
- Degraded Document Images: Scans, Faxes and Phone PhotosHow to source degraded document images for Document AI: a capture-condition taxonomy, per-page labels, fax and phone-photo traits, real vs augmented data.
- Document Dataset Requirements Template for Document AIA field-by-field requirements spec for buying document AI data: document mix, capture conditions, annotation layers, privacy, rights and acceptance tests.
- Document Fraud Detection Data: Tampered Invoices, StatementsHow to source a document fraud detection dataset: confirmed altered invoices and statements, altered-region labels, PDF file evidence and genuine controls.
- Document Layout Analysis Datasets Beyond DocLayNetCompare PubLayNet, DocBank, DocLayNet and FUNSD for layout detection, find their gaps on business documents, and specify layout data that fits production.
- Document Splitting Training Data for Scanned PacketsWhat page-stream segmentation data must contain: whole packets in original page order, first-page boundary labels, segment types and per-packet metrics.
- Document VQA Data: Questions, Answers and Answer BoxesHow to specify and source document visual question answering data: page images, natural questions, accepted answers and answer boxes from business docs.
- Engineering Drawing OCR Data: Title Blocks, Dims and GD&THow to specify and source annotated 2D engineering drawings for OCR: title block fields, dimensions, tolerances, GD&T frames, BOMs and revision blocks.
- ERP Records as Labels for Document Extraction TrainingHow to use posted ERP, TMS and claims entries as weak labels for document extraction: alignment, snapshot rules, noise audits and join keys.
- Floor Plan Datasets for Takeoff and Plan-Review AIHow to source floor plan and architectural drawing data: public set limits, labels for rooms, walls and title blocks, and commercial rights checks.
- Freight Document Extraction Data: BOLs, PODs, Rate ConsWhat a freight document extraction dataset needs: BOL, POD and rate confirmation fields, handwritten exceptions, template diversity and TMS-linked labels.
- IDP Correction Logs: Human-Validated Extraction DataHow to source human-in-the-loop document extraction data: field-level IDP correction logs, record schema, sampling bias, vendor terms and buyer checks.
- Loss Run Extraction Data: Carrier Reports With Claim LabelsWhat loss run extraction training data needs: varied carrier and TPA layouts, normalized claim labels, valuation dates, total checks and de-identification.
- Multi-Page Document Understanding Data: Long PDFsHow to specify, label and evaluate multi-page document data for long-context VLMs: cross-page tables, evidence pages, length distributions and grounding.
- Multilingual Business Document Data for OCR and ExtractionHow to source non-English and mixed-script business documents for OCR and extraction: language-per-field labels, script coverage, XFUND limits, native QA.
- Redaction Training Data: Regions, Reason Codes, DecisionsHow to source redaction training data: document regions, entity types, exemption and privilege reason codes, reviewer decisions and recall-first evals.
- Signature and Stamp Detection Data for Executed DocumentsHow to source signature, initials, stamp and seal detection data: label schema, hard negatives, public dataset limits and biometric-law checks.
- Three-Way Match Data: Linked POs, Receipts and InvoicesSource a three-way match invoice dataset: linked purchase orders, goods receipts and invoices with line-level match labels and exception resolutions.
Sources
- Pfitzmann, Auer, Dolfi, Nassar, Staar (IBM Research), "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
- CVAT.ai, "COCO (CVAT format documentation)". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
- docling-project on Hugging Face, "DocLayNet-v1.2 dataset card (README.md)". https://huggingface.co/datasets/docling-project/DocLayNet-v1.2/blob/266dc3e10783a4fa0bb528a3d841c88d68cbda9c/README.md
- Jaume, Ekenel, Thiran, "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
- Loukas et al., "EDGAR-CORPUS: Billions of Tokens Make The World Go Round" (2021). https://arxiv.org/pdf/2109.14394
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version in Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- arXiv:2111.02374, "Can I use this publicly available dataset to build commercial AI software?" (2021). https://arxiv.org/abs/2111.02374v4
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
- Microsoft (microsoft/presidio project), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- arXiv:2410.09247, "Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts" (2024). https://arxiv.org/pdf/2410.09247.pdf
- Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021; ACL 2022). https://arxiv.org/abs/2107.06499v1
- Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Klaus Krippendorff, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
- Gebru et al., "Datasheets for Datasets" (2018; Communications of the ACM 2021). https://arxiv.org/pdf/1803.09010
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.