Skip to content

Document AI data

Document AI Datasets by Task: Annotation Layers, Public Sets and Licensed Business Documents

Quick answer

Document AI datasets pair page images or PDFs with annotation layers: OCR transcriptions, layout regions, reading order, table structure, key-value fields, document classes or question-answer pairs. Choose the layer by task, then the source: public research sets such as DocLayNet and FUNSD for benchmarking and pre-training where their terms allow, synthetic pages for volume, and licensed business documents when a model must handle real issuers, poor scans, handwriting and multi-page packets.

By SourceX Editorial · Updated

Nine document AI tasks and the annotation layer each one needs

Every document AI task learns from a different annotation layer, so fix the layer, labeling unit and metric before comparing datasets or suppliers. Optical character recognition (OCR) sits underneath most rows.

TaskAnnotation layer to specifyCommon containerScore withGo deeper
Printed OCR and handwriting recognition (HTR)Text aligned to word or line boxes, with handwriting and legibility flagsALTO XML, hOCR or PAGE XML; JSON tokensCharacter and word error rate (CER, WER)OCR ground truth; handwriting data
Layout analysisRegion boxes or polygons with classes (title, text, table, picture, footnote)COCO-style JSONMean average precision across overlap thresholdsLayout datasets beyond DocLayNet
Reading orderOrdered regions or lines across columns, sidebars and form flowsOrder index per regionEdit distance to the gold sequenceReading order data
Table structure recognitionCell grid with row and column spans, header flags and cell textHTML or JSON cell listsTree-edit-distance similarity (TEDS) or cell-adjacency F1Table structure data
Key-value and line-item extractionField name, value linked to source tokens, normalized value, line-item groupingJSON with token referencesField-level precision and recall after normalizationKey-value labels; invoice line items
Classification and splittingClass per document; boundary labels inside a scanned packetPer-page CSV or JSONPer-class F1; boundary F1Classification data; page-stream segmentation
Document QA and VLM fine-tuningQuestion, answer, evidence page and region; page-to-Markdown or HTML targetsJSONLExact match or normalized string similarity; evidence accuracyDocument QA data; PDF parsing data
Charts and drawingsChart paired with its source table; title-block fields, dimensions, symbolsImages with JSON or CSVValue-level accuracy; symbol detection precision and recallChart data; engineering drawings
Redaction, signatures and stampsSensitive spans and boxes with human redaction decisions; signature, stamp and seal boxesSpan lists plus COCO-style boxesRecall per entity type; detection precisionRedaction data; signature and stamp data

Layers stack: one OCR token box is the evidence for an extracted field, a table cell's content and a question's answer region, so buy layers in one schema keyed to one coordinate frame (one annotation spec for layout, reading order, tables and fields). For layout, DocLayNet ships 80,863 manually annotated pages with boxes in 11 classes in COCO format [1]. The COCO annotation format's JSON specification includes fields for license information [2].

One page, every layer: what a deliverable record contains

A deliverable record keeps the page image, every annotation layer, each label's origin, the license reference and the privacy treatment together, so a reviewer can trace any value back to pixels and a license.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "doc_id": "pkt-20931-d02",
  "packet": {"packet_id": "pkt-20931", "pages": [3, 4], "doc_class": "vendor_invoice",
             "packet_classes": ["cover_letter", "vendor_invoice", "proof_of_delivery"]},
  "page": {"page_no": 3, "image": "pages/pkt-20931-p003.png", "dpi": 300,
           "width_px": 2550, "height_px": 3300, "rotation_deg": 0,
           "capture": "fax_bw", "box_convention": "x0_y0_x1_y1_px_top_left"},
  "ocr_tokens": [
    {"id": "t118", "text": "INV-55210", "bbox": [1890, 212, 2210, 258], "conf": 0.97, "style": "printed"},
    {"id": "t119", "text": "Net 30", "bbox": [1890, 300, 2040, 344], "conf": 0.71, "style": "handwritten"}
  ],
  "layout_regions": [
    {"id": "r2", "class": "key_value_block", "bbox": [1700, 180, 2400, 360], "reading_order": 2},
    {"id": "r6", "class": "table", "bbox": [150, 1120, 2400, 2280], "reading_order": 6}
  ],
  "table_cells": [
    {"region": "r6", "row": 2, "col": 3, "row_span": 1, "col_span": 1, "is_header": false, "tokens": ["t402"]}
  ],
  "fields": [
    {"key": "invoice_number", "value": "INV-55210", "tokens": ["t118"], "source": "erp_ap_entry"},
    {"key": "payment_terms", "value": "Net 30", "normalized": "NET_30", "tokens": ["t119"],
     "source": "annotator_double_pass"}
  ],
  "rights": {"license_ref": "LIC-0142", "third_parties": ["issuing_vendor"],
             "permitted_uses": ["model_training", "internal_evaluation"]},
  "privacy": {"actions": ["signature_masked", "remit_account_replaced_with_surrogate"],
              "layers_cleaned": ["pixels", "text_layer", "pdf_metadata"], "method_ref": "DEID-V3"},
  "split_key": "issuer:9f2c"
}

Check four things in a record like this: a declared box convention (units, corner or width-height form, origin), since mixed conventions silently misalign labels; values linked to OCR token IDs, so grounding can be scored; a source field separating system-of-record values, such as an accounts-payable entry, from annotator labels; and a split key that keeps each issuer or template on one side of a split.

Public sets, generated pages and licensed archives compared

Most document AI programs combine three routes, each failing a different check: public datasets on coverage and license clarity, synthetic pages on realism, and licensed business documents on per-deal rights and privacy work.

RouteStrong forWeak forCheck before use
Public research datasetsLayout pre-training, benchmarks, baselinesBusiness forms at volume, scan noise, line items, clear commercial termsLicense scope (pages or annotations); release; test-set overlap
Synthetic or rendered documentsVolume, rare field combinations, no personal dataIssuer and template diversity, real degradation, handwriting, stamps, packet structureAccuracy on issuers the generator never saw
Licensed business documentsReal issuers, template versions, capture conditions, operational labelsSpeed; per-deal rights review and de-identificationHolder's right to share; third-party content; redaction method

Public research datasets. Models trained on DocLayNet were more robust than those trained on PubLayNet or DocBank, which come mainly from scientific repositories [1]; a later DocLayNet v1.2 is hosted on Hugging Face by the docling-project [3], so name the release and read its card. FUNSD's 199 noisy scanned forms, annotated for text detection, OCR, layout and entity labeling and linking [4], can benchmark form understanding, but the paper notes the dataset may not be large enough to create a generalizable application. EDGAR-CORPUS, US 10-K reports from 1993 to 2020 split into items in JSON [5], helps language pre-training but teaches nothing about layout.

License metadata is unreliable. The Data Provenance Initiative's audit of more than 1,800 text datasets found license omissions above 70% and error rates above 50% on popular hosting sites [6], and a 2021 study found potential license-violation risks in five of six widely used public image datasets used commercially, partly because one dataset can mix sources under different licenses [7]. Check whether a license covers the page images or only the annotations (which public document datasets allow commercial use).

Synthetic documents. Template rendering gives volume and free labels, but a model can learn where a template puts a field instead of what the field means. Test on issuers and capture conditions the generator never produced, such as faxes, photocopies and phone photos (real-world capture conditions; where synthetic documents break).

Licensed business documents. Operational archives hold what generators rarely reproduce: many issuers, template revisions over years, margin notes, stamps over print and packets that mix document types. They often hold labels too, such as values keyed into an ERP (documents paired with system-of-record entries) and reviewers' corrections to production extraction (IDP correction logs). Compare annotating your own pages with licensing labeled ones.

SourceX sources operational datasets from US companies, including documents and finance and legal workflow records, and manages the licensing agreement. Buyers describe the documents; SourceX looks for businesses that hold them, and the supplying company approves every release. These are kinds of data it sources, not inventory under contract, so a request does not guarantee a match, and it does not source scraped public web content. See enterprise document datasets, training data for multimodal document models and the document AI and RAG use case, or describe your document requirements to SourceX.

Who else is on the page: counterparties, patients and account holders

A business document rarely belongs to one party: the company that holds it, the counterparty that issued it, and the people named in it all have interests. Check rights and privacy per document type, not per archive.

  • Counterparty terms. Contracts, supplier invoices and bills of lading name another company and its prices and terms; ask whether the holder's confidentiality clauses restrict sharing and which values to mask.
  • Health documents. CMS-1500 and UB-04 claim forms, explanations of benefits and prior-authorization faxes carry protected health information. HIPAA de-identification uses Safe Harbor, removing 18 listed identifiers of the individual and of relatives, employers or household members with no actual knowledge of identifiability, or an expert determination that identification risk is very small [8]. For health records, SourceX requires one of these methods before anything is considered for a license.
  • Financial documents. Bank statements, pay stubs and loan files held by lenders carry nonpublic personal information. Under Regulation P (12 CFR 1016.11), a recipient that gets it from a nonaffiliated financial institution under an exception may use and disclose it only in the ordinary course of business to carry out the purpose for which it was received [9], so ask how the holder obtained it and what its own privacy notice allows it to share (bank statements and income documents).
  • Pixels, text layers and metadata. A PDF holds personal data in rendered pixels, the text or OCR layer and metadata; masking one leaves the others readable. Presidio, an open-source PII detection and anonymization SDK, states there is no guarantee it finds all sensitive information [10] (redacting PII in scanned documents). Black boxes also teach a model black boxes, so ask which values were replaced in place with realistic surrogates.
  • Published PDFs for pre-training. For general-purpose models placed on the EU market, Article 53 of the AI Act requires a copyright policy that honors text-and-data-mining reservations and a public summary of training content [11] (licensed PDF corpora for pre-training).

On deals SourceX manages, every dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place; the license defines included records, allowed uses, term and delivery. Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded per dataset and a sample is checked after processing; no de-identification method is perfect. Compare methods in the de-identified data buyer's guide.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Scoring document models on documents they have never seen

Evaluate document models on a private held-out set split by issuer, template and time and scored at field level, because public benchmarks may already sit in pre-training data and repeated templates inflate scores.

Contamination is the first reason. A 2024 study notes that many large language models' training data is contaminated with test data, and that private holdout sets that could verify scores do not exist for most benchmarks [12]. Assume a document VLM may have seen public document benchmarks.

Template repetition is the second: one vendor can send thousands of invoices from one template. Language-model corpora show the same pattern: they contain many near-duplicate examples, and models trained on deduplicated versions emitted memorized text about ten times less often [13]. Deduplicate near-identical pages, then keep each issuer and template family on one side of the split.

Gold labels need an audit too. Test sets of 10 widely used vision, language and audio datasets had an estimated average label error rate of at least 3.3%, and such errors can change which model ranks best [14]. Double-annotate a subset, as DocLayNet did [1], and report Krippendorff's alpha, where 1 is perfect reliability and 0 is none beyond chance [15].

Write scoring rules into the specification: date, currency and amount normalization; scoring for fields absent from the page; line-item matching when rows split or merge; and whether a correct value with the wrong evidence box counts. See document extraction evaluation sets, private evaluation sets vs public benchmarks and evaluation datasets from real business work.

Start here: business document types and their guides

Each record type links to its cluster guides and to the SourceX page on licensing it.

Record typeCluster guidesSourceX page
Invoices, purchase orders, receiptsInvoice line items; receipts; three-way matchInvoices and receipts; purchase orders
Scanned and handwritten formsACORD forms; checkbox and selection marksScanned forms and handwriting; underwriting files
Financial and bank statementsFinancial statement spreading; bank statementsFinance and accounting AI data
Contracts and agreementsContract clause annotation; contract familiesContract redlines
Insurance and health claimsLoss run extractionInsurance claims; healthcare revenue cycle
Freight and trade documentsFreight documents; trade documents; trade financeSupply chain and logistics
Drawings and diagramsFloor plans; P&IDsCAD and PCB design files
Quality and safety documentsCertificates of analysis; safety data sheetsManufacturing quality records
Mixed archives and long PDFsLicensed PDF corpora; long documents; multilingual documentsEnterprise document archives

Also see document fraud detection data, document version pairs, the multimodal training data hub and the AI data buyer's guide.

Eight lines to put in a document data request

A document data request is complete when a supplier can tell from it which pages qualify, which layers to deliver and how you will accept them.

  1. Tasks and layers. The tasks from the first table and their layers, in one schema.
  2. Document mix. Record types, issuers or templates per type, pages per document, and whether packets stay intact.
  3. Capture conditions. Born-digital PDFs versus scans, faxes and phone photos; resolution; share of pages with handwriting, stamps or signatures.
  4. Languages and scripts. Including mixed-script and right-to-left pages.
  5. Label origin and quality. System-of-record values or annotator passes, agreement targets and adjudication.
  6. Splits and holdouts. Issuer- or template-disjoint splits and an evaluation slice that never enters training.
  7. Privacy and rights. Fields masked or replaced across pixels, text layers and metadata, and the uses you need licensed.
  8. Documentation and delivery. A datasheet covering motivation, composition, collection process and recommended uses [16], plus formats and transfer (delivery formats and transfer).

The document dataset requirements spec turns this list into a full request.

Looking for real business documents for document AI?

Describe the document types, annotation layers, volumes, capture conditions and allowed uses you need on the SourceX buyer page. SourceX looks for US companies that hold those documents, checks the data and the supplier's licensing permissions, manages the license and coordinates delivery and payment; nothing is contracted until a supplier agrees. Describe the documents your model needs.

Guides in this section

Sources

  1. Pfitzmann, Auer, Dolfi, Nassar, Staar (IBM Research), "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
  2. CVAT.ai, "COCO (CVAT format documentation)". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
  3. docling-project on Hugging Face, "DocLayNet-v1.2 dataset card (README.md)". https://huggingface.co/datasets/docling-project/DocLayNet-v1.2/blob/266dc3e10783a4fa0bb528a3d841c88d68cbda9c/README.md
  4. Jaume, Ekenel, Thiran, "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
  5. Loukas et al., "EDGAR-CORPUS: Billions of Tokens Make The World Go Round" (2021). https://arxiv.org/pdf/2109.14394
  6. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version in Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
  7. arXiv:2111.02374, "Can I use this publicly available dataset to build commercial AI software?" (2021). https://arxiv.org/abs/2111.02374v4
  8. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  9. Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
  10. Microsoft (microsoft/presidio project), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  11. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  12. arXiv:2410.09247, "Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts" (2024). https://arxiv.org/pdf/2410.09247.pdf
  13. Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021; ACL 2022). https://arxiv.org/abs/2107.06499v1
  14. Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  15. Klaus Krippendorff, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  16. Gebru et al., "Datasheets for Datasets" (2018; Communications of the ACM 2021). https://arxiv.org/pdf/1803.09010

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data