Skip to content

Evaluation and benchmarking datasets

Document extraction evaluation sets: field-level ground truth from real forms and invoices

Quick answer

A document extraction evaluation dataset pairs real document images or PDFs with the correct value for every target field, so you can score an IDP pipeline or vision-language model field by field. The strongest ground truth for business documents is often the value a company actually posted in its system of record (the ERP invoice header, the AP line, the claim record), reconciled against the page. Buyers should specify the field schema, normalization rules, null handling, layout slices and de-identification method before any documents change hands.

By SourceX Editorial · Updated

Why public KIE benchmarks are not enough for an invoice extraction accuracy test

Public benchmarks are useful for comparing architectures but rarely match your vendors, field schema or scan conditions. DocILE defines key information localization and extraction (KILE) and line item recognition (LIR) tasks over annotated business documents [1], RealKIE packages five enterprise sets such as SEC S-1 filings, NDAs and FCC invoices [2], and FUNSD covers 199 noisy scanned forms [3]. Those are good starting points, and recent invoice-extraction work still evaluates against the KILE and LIR conventions while building small custom annotated sets of its own [4].

The gap shows up in three places. First, the schema: your downstream system may need remit_to_address, po_number and tax_amount split by jurisdiction, which no public set labels. Second, the distribution: a model that scores well on FUNSD scans may fail on your top 40 suppliers' templates. Third, contamination: widely mirrored public documents may already sit in a foundation model's pretraining data, which is why held-out private sets matter (see private eval sets vs public benchmarks and contamination-resistant evaluation design). Always check each public set's license before using it for commercial model selection.

System-of-record values vs fresh human labels as ground truth

System-of-record values are the cheaper and often more faithful truth for "what the business needed," while fresh human labels are better for "what is literally printed on the page." An invoice posted in an ERP such as SAP S/4HANA, Oracle NetSuite or Microsoft Dynamics records the vendor, invoice number, dates, totals and GL lines that an AP clerk accepted, often after a three-way match against the PO and receiving record. That value reflects business reality, but it can diverge from the printed document.

Typical divergences you must handle explicitly:

  • Corrections at posting. A clerk fixes a wrong tax total or overrides a due date per contract terms; the posted value is "right" for the business but not printed on the page.
  • Master-data substitution. The ERP stores a vendor ID and canonical name, not the letterhead spelling.
  • Aggregation. Ten printed line items may be posted as two GL lines.
  • Credits and partial postings. One document can map to several records, or a record to several documents.

The practical answer is a hybrid: use posted values as candidate truth, then have reviewers confirm each field against the page and tag it printed_match, corrected_at_posting or not_on_page. Even curated test sets carry errors; Northcutt et al. estimate an average label error rate of at least 3.3% across test sets of 10 widely used datasets [5]. Double-annotating a subset, as DocLayNet did to measure inter-annotator agreement [6], gives you a ceiling to interpret model scores against. For a broader method, see building a golden evaluation dataset from business records.

Field-level metrics, normalization rules and null handling

Field-level evaluation needs a written scoring contract, because the same output can score 70% or 95% depending on normalization. Define per field: the comparison type (exact, normalized exact, numeric tolerance, fuzzy string), the normalization steps, and how empty values score.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldTypeNormalization before comparisonMatch ruleNull rule
invoice_numberstringstrip whitespace, uppercase, drop leading "INV-" only if schema says soexactabsent on page = expected null
invoice_datedateparse to ISO 8601 (YYYY-MM-DD); resolve MM/DD vs DD/MM by vendor localeexactnull only if not printed
total_amountdecimalremove currency symbols and thousands separators; 2 decimalsabs diff <= 0.01never null on a valid invoice
currencycodemap symbols to ISO 4217exactinfer only if schema allows
vendor_tax_idstringremove spaces and hyphensexactexpected null if not printed
line_items[]arraymatch rows by description + amount alignmentper-cell F1, row-level exactempty array if none

Report precision, recall and F1 per field, not only document-level "straight-through" rates, and separate three null outcomes: correct abstention (expected null, predicted null), hallucination (expected null, predicted value) and miss (expected value, predicted null). Hallucinated values on optional fields are often the costliest failure in AP automation because they pass silently. Line items need their own metric: align predicted rows to gold rows first, then score cells, as LIR-style tasks do [1].

Add schema validity as a gate before accuracy. Teams increasingly treat a failed structured output, such as malformed JSON or a missing required key, as an evaluation signal in its own right [7]; validate every prediction against a JSON Schema and count invalid outputs as failures for every field they contain.

Layout diversity slices: vendors, scans and handwriting

A useful extraction eval set is stratified so that a single aggregate score cannot hide failure on the documents that matter. Slice dimensions to request and record per document:

  • Template family: vendor or issuer, with a cap per vendor so the top three suppliers do not dominate.
  • Capture channel: born-digital PDF, scanned at 200 or 300 dpi, fax, mobile photo with skew and shadow.
  • Content features: multi-page invoices, line-item tables that break across pages, stamps, handwritten annotations, rotated pages.
  • Language and locale: date formats, decimal commas, multi-currency documents.
  • Rare but costly cases: credit memos, duplicate submissions, invoices with both a PO and non-PO section.

Allocate examples deliberately to rare, high-risk cells rather than sampling proportionally; stratified evaluation sets covers sizing. If your pipeline includes a separate OCR stage, pair this set with OCR ground truth aligned to real scans so you can attribute errors to recognition vs field mapping. Annotation conventions for training labels are covered separately in key-value extraction labels.

Request template for an extraction evaluation set

The fastest way to get a usable set is to describe the documents, the truth source and the scoring contract up front. Suppliers can only say whether they hold matching records if the request is specific.

Illustrative example: invented to show structure; it does not describe an available dataset.

eval_set_request:
  task: field-level extraction (header fields + line items)
  document_types: [vendor invoices, credit memos, remittance advices]
  capture_mix: {born_digital_pdf: 50%, scanned_300dpi: 35%, mobile_photo: 15%}
  slices: {max_docs_per_vendor: 15, min_handwritten_annotations: 60, multipage_tables: 100}
  ground_truth_source: ERP posted values, reconciled to page by reviewer
  per_field_provenance: [printed_match, corrected_at_posting, not_on_page]
  schema: invoice_v3.json   # JSON Schema with required/optional fields
  normalization_rules: scoring_contract_v2.md
  double_annotated_subset: 10%
  deidentification: replace person names, emails, phones, bank account numbers; preserve layout geometry
  delivery: page images (TIFF/PNG) + original PDFs + gold JSON + bounding boxes where available
  holdout: never used for training or prompt examples

Keep the evaluation set sealed: log who accessed it, never paste its documents into prompts during development, and rotate a fresh slice in periodically so the score stays meaningful.

De-identifying personal fields while preserving layout

De-identification for extraction eval must remove personal data without destroying the geometry the model is being tested on. Blacking out a payee name changes the visual signal; replacing it with a synthetic value of similar length and font, in the same bounding box, preserves layout difficulty. The replacement must be applied consistently in the pixels, the OCR or text layer, PDF metadata, and the gold JSON, or your ground truth will no longer match the page. Scanned documents carry extra traps, covered in redacting PII in scanned documents.

Health documents such as EOBs, claim forms and referral letters fall under HIPAA, where de-identification follows Safe Harbor or Expert Determination and limited data sets carry their own rules [8]. If your extraction model feeds a system classed as high-risk under the EU AI Act, Article 10 sets quality and governance criteria for training, validation and testing data sets alike [9]; Regulation (EU) 2026/1744 amended Article 10 and moved high-risk application dates to December 2027 and August 2028 [9]. This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

How SourceX sources extraction evaluation data

SourceX sources operational datasets from US companies, including documents and finance and legal workflow records, and manages the commercial process through licensing. Data is sourced on request rather than held in stock, so a request does not guarantee a match; buyers describe the documents and field truth they need, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees.

Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect; health records require HIPAA de-identification. Diligence materials are prepared per dataset, and delivery runs through private, access-controlled workflows only after an executed agreement. SourceX does not train models and does not publish prices. You can start a request on the SourceX buyers page. For training rather than evaluation data, see training data for document understanding models, document understanding use cases, licensed invoices and receipts, scanned forms and handwritten documents and evaluation sets built from real business work. The wider map of eval data lives in the evaluation datasets hub.

Source a document extraction evaluation set with verified field values

SourceX sources documents and finance workflow records from US companies on request, with every dataset rights-reviewed and delivered under a license that defines records, uses, term and delivery. Describe the document types, field schema and ground-truth source you need at sourcex.si/buyers.

Frequently asked questions

How many documents does a document extraction evaluation set need?

Size follows the slices, not a global number. Decide the smallest per-slice count that gives a confidence interval narrow enough to separate the models you are comparing, then multiply across vendor, capture and content slices you must report on.

Can posted ERP values be used without any human review?

Not safely. Posted values encode corrections, master-data substitutions and aggregations, so an unreviewed set will penalize models for correctly reading what is printed. A reviewer pass tagging each field's provenance makes both "page truth" and "business truth" scoring possible.

Should the eval set include line items or only header fields?

Include line items if your workflow posts them. Line-item tables are where multi-page breaks, merged cells and row alignment fail, and header-only scores routinely overstate production accuracy.

Sources

  1. arXiv (Šimsa et al.), "DocILE Benchmark for Document Information Localization and Extraction" (2023). https://arxiv.org/pdf/2302.05658
  2. arXiv, "RealKIE: Five Novel Datasets for Enterprise Key Information Extraction" (2024). https://arxiv.org/pdf/2403.20101
  3. arXiv (Jaume, Ekenel, Thiran), "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
  4. arXiv, "Invoice Information Extraction: Methods and Performance Evaluation" (2025). https://arxiv.org/pdf/2510.15727
  5. arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  6. arXiv (IBM Research), "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
  7. OneUptime, "How to Build a Golden Evaluation Dataset from Real LLM Production Failures" (2026). https://oneuptime.com/blog/post/2026-08-31-build-golden-evaluation-dataset-production-failures/markdown
  8. eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  9. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data