Document AI data
Key-Value Extraction Labels: Annotated Forms and Semi-Structured Business Documents
Quick answer
A usable key information extraction (KIE) dataset gives every target field a typed value, the value's location on the page, the printed key it answers to, and a normalized form you can score against. Public sets such as FUNSD, CORD, SROIE and DocILE are good starting benchmarks, but they cover narrow domains or a fixed set of issuers. Production IDP teams usually need licensed business documents labeled to their own field schema, with template diversity tracked and held out for testing.
By SourceX Editorial · Updated
What the public KIE benchmarks cover, and where they stop
Public KIE sets are useful for architecture choices and smoke tests, but each one fixes a domain, a label schema and a license you have to check before commercial use. FUNSD has 199 scanned forms labeled with entity types (question, answer, header, other) and the links between them [1]. That makes it the reference for key-value linking, and also a reminder of scale: 199 pages cannot represent the template variety of a real intake queue.
DocILE is the largest business-document benchmark in the group, with about 6.7k annotated business documents and roughly 100k synthetic ones, covering both key information localization and line-item recognition [2]. RealKIE adds five enterprise collections (SEC S1 filings, NDAs, UK charity reports, FCC invoices, resource contracts) and records the conditions that break models in practice: poor text serialization, sparse labels across long documents and complex tables [3]. Kleister NDA (540 documents) and Kleister Charity (2,788 reports, 61,643 pages) test extraction from long formal documents, where the answer is often not next to a printed key [4].
A toolkit dataset list from PaddleOCR indexes public sets including FUNSD, the multilingual XFUND and the WildReceipt receipt set; MMOCR's dataset docs cover WildReceipt [5][6]. Treat those lists as an index, not a license review: confirm each dataset's terms, size and annotation format against its primary source before any commercial training run. A published critique of business-document benchmarks argues that common problem definitions and benchmarks do not reflect domain-specific aspects and practical needs of business document extraction [7], and multi-task sets such as BuDDIE show the value of labeling several tasks on the same pages [8].
Label granularity: the layers a KIE annotation should carry
Specify labels as layers, because each layer supports a different model family and evaluation. A token classifier in the LayoutLM family needs word boxes and BIO tags; a vision-language model fine-tuned for structured output needs page images paired with a JSON target; an evaluation harness needs normalized values. If you buy only one layer, you will rebuild the others by hand.
The core layers are:
- Value text and value box. The exact string as printed, plus a bounding box or polygon in page coordinates, with the page index for multi-page files.
- Key text and key box. The printed label ("Invoice No.", "Policy Period", "Ship To"), which may be absent, abbreviated or in another language.
- Key-value link. An explicit edge between the key entity and the value entity, the structure FUNSD formalized [1]. Links matter when one key governs several values or a value sits far from its key.
- Field type. Your schema name (invoice_number, due_date, remit_to_address), independent of the printed key wording.
- Normalized value. ISO 8601 dates, decimal amounts with an ISO 4217 currency code, and canonical identifiers, so "03/04/25" is resolved to a single date by locale.
- Absence and ambiguity flags. A field marked "not present" is a label; an empty cell is not. Flag illegible, crossed-out or conflicting values instead of dropping them.
Template diversity matters more than page count
For field extraction, the number of distinct issuers and layouts predicts generalization better than raw page volume. Ten thousand invoices from 40 vendors teach a model 40 layouts; two thousand from 900 vendors teach it to read. Ask suppliers to report issuer counts, pages per issuer and the long tail, not just totals.
Split by template group, never by page. If pages from the same issuer land in both train and test, field scores will look strong and collapse on the first new vendor. Hold out a set of issuers entirely, and keep a second holdout of document types the model has never seen if you plan to claim cross-type generalization. The document extraction evaluation ground truth guide covers how to freeze and version that test set.
Real documents also carry conditions that synthetic generators miss: stamps over fields, handwritten corrections, fax artifacts, rotated scans and continuation pages. The comparison of synthetic and real documents explains where generated layouts stop transferring.
Using system-of-record values as labels, and aligning them back to the page
Normalized values already stored in an ERP, AP or policy-admin system can serve as extraction labels, but only after they are aligned to locations on the page. An AP system holds the posted invoice total, not the box where it was printed, and the posted value may have been corrected, converted or split across lines. Unaligned values train a model to guess, and they cannot score localization.
A practical alignment pass searches OCR tokens for candidate matches to each stored value, accepts exact or normalized matches, and routes multi-match or no-match fields to human review. Record the match method per field so you can filter weak labels later. The documents paired with ERP records guide covers that pairing in depth; this page covers the annotation layer you align to.
Specification template for a KIE data request
A good request names the field schema, label layers, diversity targets and evaluation splits before anyone quotes. The template below is the minimum a supplier or annotation vendor needs to respond.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Spec item | Example entry |
|---|---|
| Document types | Vendor invoices, purchase orders, bills of lading, W-9s |
| Field schema | 24 header fields per type, with types (string, date, money, address, ID) and a "not present" rule |
| Label layers | Value text and box; key text and box; key-value links; field type; normalized value; ambiguity flag |
| Coordinate convention | Pixel coordinates at stated DPI, top-left origin, page index per box |
| Diversity target | Issuer count per type, cap on pages per issuer, mix of native PDF and scanned pages |
| Splits | Train, validation, and a test set of issuers held out entirely |
| Format | Page images (PNG or TIFF) plus JSON per page; optional FUNSD-style entity and link lists |
| Quality evidence | Double-annotated sample with field-level agreement; adjudication notes |
| Redaction | Personal identifiers removed or replaced, with the method stated and spans flagged |
| Rights | Supplier confirms ownership and that use for model training is permitted under the license |
An illustrative per-page record might look like this:
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"doc_id": "inv-000412",
"page": 1,
"issuer_group": "vendor-0187",
"fields": [
{
"field": "invoice_date",
"key": {"text": "Inv. Date", "box": [112, 208, 190, 226]},
"value": {"text": "03/04/25", "box": [204, 208, 288, 226]},
"link": true,
"normalized": "2025-03-04",
"flags": []
},
{
"field": "total_amount",
"key": {"text": "Amount Due", "box": [980, 1544, 1110, 1562]},
"value": {"text": "$12,480.00", "box": [1140, 1544, 1262, 1562]},
"link": true,
"normalized": {"amount": "12480.00", "currency": "USD"},
"flags": []
},
{"field": "po_number", "status": "not_present"}
]
}
Evaluating extraction: per-field metrics by template group
Report per-field precision, recall and exact match on normalized values, broken out by template group and document type. A single micro-averaged F1 hides the fields that matter commercially, such as totals, due dates and remit-to accounts, which are rare compared with line descriptions. Score localization separately (box overlap with the labeled value) if your product highlights evidence to reviewers.
Common failure modes to test for:
- Key-value swaps on two-column forms, where the model reads the adjacent column's value.
- Locale errors on day-month order and decimal separators, visible only with normalized labels; see multilingual business document data.
- Long-document misses where the field appears once on page 14, the pattern RealKIE and Kleister document [3][4].
- Table spillover when header fields sit inside or next to a line-item table; the invoice line-item extraction guide covers the table layer.
- Hallucinated values from generative extractors when a field is absent, which only a "not present" label can catch.
Rights and privacy checks before KIE data reaches training
Business forms carry personal and third-party data in exactly the fields KIE targets: names, addresses, account numbers and tax IDs. Decide whether you need real values for those fields or consistent surrogates, and require the supplier to state the redaction or replacement method and which spans were changed. Replacement that preserves format (a plausible account number in the same box) keeps the extraction task intact.
Confirm who owns the documents and whether the issuer's content raises confidentiality questions; the guides on copyright in business records and third-party confidential information screening cover those checks. If you plan to annotate in-house rather than license labeled pages, weigh the cost using annotating your own documents vs licensing pre-labeled ones.
How SourceX approaches labeled business document requests
SourceX sources operational datasets, including documents and finance and legal workflow records, from US companies on request; it does not hold them in stock, and a request does not guarantee a match. Buyers describe the data they need, such as the field schema and diversity targets above, and SourceX looks for US businesses that hold it. Every release is approved by the supplying company, rights-reviewed for ownership and consents, and delivered under a license that defines records, uses, term and delivery.
Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Diligence materials on source, rights, preparation and allowed use are prepared per dataset. Start a request on the SourceX buyers page, and browse related record types such as invoices and receipts, purchase orders, scanned forms and handwritten documents and expert annotations and labels. For background on entity tagging, see the named entity recognition glossary entry, and for the full set of document tasks, the Document AI data hub and AI data hub.
Request key information extraction data from real business documents
If public KIE benchmarks are too small or too narrow for your field schema, describe the document types, fields, label layers and issuer diversity you need. SourceX looks for US businesses that hold matching documents, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Nothing is contracted until a supplier agrees. Describe your KIE data needs to SourceX.
Sources
- arXiv (Jaume, Ekenel, Thiran), "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
- arXiv (Šimsa et al.), "DocILE Benchmark for Document Information Localization and Extraction" (2023). https://arxiv.org/pdf/2302.05658
- arXiv, "RealKIE: Five Novel Datasets for Enterprise Key Information Extraction" (2024). https://arxiv.org/pdf/2403.20101
- arXiv, "Kleister: Key Information Extraction Datasets Involving Long Documents with Complex Layouts" (2021). https://arxiv.org/abs/2105.05796v1
- PaddleOCR documentation, "Key Information Extraction datasets". https://www.paddleocr.ai/v3.3.2/en/datasets/kie_datasets.html
- MMOCR documentation, "Key Information Extraction". https://mmocr.readthedocs.io/en/v0.3.0/datasets/kie.html
- arXiv, "Business Document Information Extraction: Towards Practical Benchmarks" (2022). https://arxiv.org/pdf/2206.11229
- arXiv, "BuDDIE: A Business Document Dataset for Multi-task Information Extraction" (2024). https://arxiv.org/pdf/2404.04003
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.