Document AI data
Invoice Line-Item Extraction Data Across Vendors, Templates and Years
Quick answer
An invoice line-item extraction dataset is useful when it spans many vendors and template versions, labels every line-table row (description, quantity, unit price, amount, tax, SKU, PO line reference), handles multi-page tables, and ties labels to what accounts payable actually posted. Public sets such as DocILE give a benchmark structure but cover a limited slice of real AP variety. Production gains usually come from licensed operational invoices with vendor-level held-out splits and documented labeling.
By SourceX Editorial · Updated
Why vendor and template count matter more than invoice count
Generalization in invoice extraction tracks the number of distinct layouts a model sees, not the raw number of pages. Ten thousand invoices from 40 vendors teach 40 line-table geometries; the same volume from 3,000 vendors teaches column orders, merged description cells, discount rows and tax subtotals your model will meet in production. Treat this as a working hypothesis to test on your own data, but it matches how DocILE was built: its authors group documents by layout cluster and keep unseen layouts in the test split [1].
Ask suppliers for a vendor histogram, not just a page count. A healthy AP corpus has a long tail: a few high-volume vendors (utilities, freight carriers, SaaS subscriptions) and hundreds of vendors that send a handful of invoices per year. That tail is where line-item recall collapses, so it is the part worth paying for.
Years matter too. Vendors redesign templates, switch billing systems and add QR codes or remittance stubs, so a corpus spanning several fiscal years carries template drift that a single year cannot.
What public invoice datasets cover and where they stop
The main public reference for line items is DocILE, which defines key information localization and extraction (KILE) and line item recognition (LIR) tasks over about 6.7k annotated business documents plus roughly 100k synthetic documents [1]. A 2025 evaluation paper follows the same KILE and LIR framing but had to build its own 102-invoice ground truth, mixing born-digital and scanned files, to measure line items [2]. That gap is telling: beyond DocILE, no large, widely used public set of real invoices with complete line-item ground truth has emerged.
Practitioner guides make the same point from the product side. Public sets rarely prove production accuracy, and synthetic invoices miss the stamps, skew, handwritten approvals and broken tables that real AP inboxes contain [3]. Use DocILE to pick metrics and annotation conventions, then source real documents for training and evaluation. Broader options are mapped on the Document AI datasets hub.
Line-item and header labels a usable set needs
A usable label set records header fields once per invoice and line-item fields once per row, with bounding boxes and a stable row ID. Strong invoice datasets ship bounding boxes, line-item rows, metadata and clean train, validation and test splits [3]. Without row IDs, you cannot score row-level matching, and without boxes you cannot train layout-aware models such as the LayoutLM family or score field localization.
Header fields typically include vendor name and address, vendor tax ID, invoice number, invoice date, due date, PO number, currency, payment terms, subtotal, tax total, freight and invoice total. Line fields include description, quantity, unit of measure, unit price, line amount, tax rate or amount, SKU or vendor part number, and PO line reference. Also label non-item rows (discounts, surcharges, deposits, subtotals) with a row type, because misreading a subtotal as an item is a common failure.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"doc_id": "inv_000412",
"vendor_key": "V-1187",
"template_cluster": "T-1187-b",
"source_type": "scanned_pdf",
"pages": 3,
"header": {
"invoice_number": {"value": "88-20417", "page": 1, "bbox": [412, 96, 520, 112]},
"invoice_date": {"value": "2025-03-14", "page": 1},
"po_number": {"value": "PO-55102", "page": 1},
"currency": "USD",
"subtotal": "4,812.50", "tax_total": "385.00", "invoice_total": "5,197.50"
},
"line_items": [
{"row_id": 7, "row_type": "item", "page": 2, "continued_from_page": 1,
"description": "Hex bolt M10x40 zinc", "sku": "HB-1040Z",
"quantity": "500", "uom": "EA", "unit_price": "0.42", "amount": "210.00",
"tax_rate": "8.0", "po_line_ref": "PO-55102/3"},
{"row_id": 8, "row_type": "discount", "page": 2, "amount": "-21.00"}
],
"posted_ap": {"gl_account": "5120", "cost_center": "PLANT-02", "posted_total": "5,197.50"},
"masking": {"vendor_bank_account": "tokenized", "remit_to_contact": "removed"},
"split": "test_heldout_vendor"
}
Multi-page line tables and the failure modes they expose
Multi-page invoices are where line-item models fail most visibly, so the dataset should flag and label table continuation explicitly. Typical breaks include header rows repeated on each page, a row whose description wraps across a page break, carried-forward subtotals ("balance forward") that look like items, and totals that appear only on the last page.
Ask for a continued_from_page marker or equivalent, page-level boxes, and a check that the sum of line amounts reconciles to the stated subtotal. Reconciliation failures are a cheap way to find both extraction errors and label errors. Label noise is not a niche concern: an audit of 10 widely used benchmark test sets estimated an average label error rate of at least 3.3% [4], so a reconciliation pass on your gold set is worth the effort.
Using posted AP records and structured e-invoices as labels
The highest-value labels often come from systems of record: the AP entry that was actually posted, with vendor ID, amounts, GL account and cost center. Posted records label header fields and coding models directly, though they do not always map one-to-one to printed lines (one line can be split across GL accounts, or several lines rolled up). The pairing approach is covered in documents paired with system-of-record entries, and full PO, receipt and invoice chains are covered on three-way match data.
Structured e-invoices offer exact labels where they exist. Hybrid formats such as Factur-X and ZUGFeRD embed UN/CEFACT Cross Industry Invoice XML, aligned to the EN 16931 semantic model, inside a PDF/A-3 that also renders visually; confirm profiles and validation rules against the official FeRD and FNFE-MPE specifications. Pairing the rendered PDF with its XML gives line-level ground truth without manual annotation, though layouts from e-invoicing tools are cleaner than paper invoices, so mix them with scanned and emailed PDFs rather than relying on them alone.
Human correction logs from production extraction systems are a third label source. They capture exactly where a model went wrong; see IDP human correction data.
Building an invoice extraction evaluation set you can trust
An evaluation set should hold out entire vendors and template clusters, not random pages, or scores will overstate production accuracy. Random page splits leak vendor layouts into test, which is a common reason offline line-item F1 does not survive deployment.
Practical rules for the set:
- Hold out 10–20% of vendors entirely, including some long-tail vendors with fewer than five invoices.
- Score header fields by exact or normalized match, and line items by row-level matching (all fields in a row correct) as well as field-level F1.
- Double-annotate a subset and measure agreement before freezing the gold set.
- Keep the evaluation set out of any vendor or annotator pipeline that also produces training data.
Field-level ground truth design is covered in depth on document extraction evaluation sets.
Request checklist for sourcing invoice line-item data
A precise request names the label schema, the variety you need and the masking you will accept. Use this as a starting point.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Requirement | What to specify | Why it matters |
|---|---|---|
| Vendor diversity | Minimum distinct vendors; long-tail share; vendor histogram | Drives layout generalization |
| Time span | Fiscal years covered; known billing-system changes | Captures template drift |
| Source types | Born-digital PDF, scanned PDF, email-body, e-invoice XML share | Matches your production mix |
| Header labels | Field list with boxes and normalized values | Trains and scores header extraction |
| Line labels | Row ID, row type, item fields, PO line reference | Enables row-level scoring |
| Multi-page | Share of invoices over two pages; continuation markers | Exposes table-break failures |
| System-of-record link | Posted AP entry, GL account, PO link | Labels coding and match models |
| Masking | Vendor bank details, remit-to contacts, signatures | Third-party data handling |
| Documentation | Data card with sources, labeling method, known gaps [5] | Supports internal review |
| Splits | Vendor-level held-out test set | Honest evaluation |
Third-party data inside licensed invoices
Invoices carry data about the vendors who issued them, so masking must be agreed before delivery, not after. Vendor bank account and routing numbers, remit-to contacts, signatures and sometimes employee names on ship-to lines are third-party data inside the buyer's licensed corpus. Decide whether vendor names stay (useful for vendor-normalization models) or become stable pseudonymous keys, and whether bank fields are tokenized or removed.
SourceX removes or replaces personal details such as names, emails, phone numbers and account numbers before delivery, records the method used and checks a sample; no method is perfect, so plan your own spot checks. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery. For broader invoice and receipt licensing, see license invoices and receipts for AI training; for disputed or blocked invoices, see invoice exception records.
How SourceX sources invoice line-item data
SourceX sources operational datasets, including finance workflows such as AP invoices, from US companies and manages the commercial process from licensing through ongoing purchases. Data is sourced on request rather than held in stock, so a request does not assure a match; you describe the data, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. The process runs Find, Assess, Agree, Transact and Manage, and delivery happens through private, access-controlled workflows only after an executed agreement. You can describe your invoice data requirements to SourceX using the checklist above.
Request invoice line-item extraction data
Describe the vendors, templates, years, labels and masking you need, and SourceX will look for US companies that hold matching AP invoices and manage licensing if a supplier agrees. Diligence materials on source, rights, preparation and allowed use are prepared per dataset. Start a request at sourcex.si/buyers.
Sources
- Šimsa et al., "DocILE Benchmark for Document Information Localization and Extraction" (2023). https://arxiv.org/pdf/2302.05658
- arXiv, "Invoice Information Extraction: Methods and Performance Evaluation" (2025). https://arxiv.org/pdf/2510.15727
- Invoice Data Extraction, "Invoice dataset guide". https://invoicedataextraction.com/blog/invoice-dataset
- Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Pushkarna, Zaldivar, Kjartansson, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.