Document AI data
Three-Way Match Data: Linked Purchase Orders, Goods Receipts and Invoices
Quick answer
A usable three-way match dataset is not a folder of invoices. It is a linked set: the purchase order, the goods receipt or receiving record, and the supplier invoice for the same transaction, joined by PO number and line references, with a line-level match status, the tolerance policy that produced it, and, for exceptions, how a person resolved them. Public data covers the event flow but not the documents, so teams building matching models usually need licensed real accounts-payable records.
By SourceX Editorial · Updated
What a linked three-way match record must contain
A three-way match record is useful only when every document in it can be tied back to the same PO line and the outcome is recorded at that line, not at the invoice header. Header-level "matched / not matched" flags teach a model almost nothing about where a quantity or price variance came from. Ask for the three document images (or native PDFs) plus the structured ERP rows that were used to post them.
The minimum keys are the PO number and PO line, the goods receipt document and line, the invoice number and invoice line, and the vendor identifier in a masked but consistent form. Around those keys sit the values that drive matching: ordered, received and invoiced quantity; unit of measure; unit price on the PO and on the invoice; currency; tax and freight lines; and posting dates. In SAP-style systems the history of edits to a PO or invoice lives in change documents, a header table (CDHDR) and an item table (CDPOS) holding old and new values [7], and that history is often where the real resolution story is.
Owner pages for the individual document types cover them on their own: see purchase orders for AI training and invoices and receipts. This page is about the joined set and its labels.
Why public three-way match data falls short
Public sources describe the matching process but rarely include the matched documents themselves. The best-known real example is the BPI Challenge 2019 log from a multinational coatings and paints company, in which each purchase order item carries a matching category such as three-way match with invoice after goods receipt, three-way match with invoice before goods receipt, two-way match, or consignment [1]. An object-centric conversion of that log is also published [2], and the OCEL 2.0 format lets one event reference an order, a receipt and an invoice at once [3].
Those logs are event sequences, not document images, so they cannot train extraction or cross-document matching models on real layouts. Simulated procure-to-pay logs exist too [4], but they inherit the simulator's assumptions about variance rates and resolution paths. On the document side, DocILE established key information extraction (KILE) and line item recognition (LIR) tasks for business documents [5], and a 2025 invoice extraction study still had to build its own small annotated set [6]. Neither links an invoice to its PO and receipt.
Searching for "three-way match dataset" also returns database three-way joins and three-mode data arrays, so buyers should phrase requests as "purchase order invoice matching dataset" or "PO, receipt and invoice matching data" when briefing suppliers or searching.
Labels that carry the most training value
The highest-value labels are exception outcomes with resolution notes, because clean matches are cheap to generate and abundant. A model that only sees auto-matched lines learns the easy case. What finance-automation teams need are the lines that failed tolerance and what happened next.
Useful exception classes include:
- Price variance: invoice unit price outside the PO price tolerance (percentage, absolute amount, or both).
- Quantity variance: invoiced quantity exceeds received quantity, or receipt exceeds order.
- Missing receipt: invoice arrives before the goods receipt is posted; often resolved by waiting, not by override.
- Unit of measure mismatch: cases on the PO, eaches on the invoice.
- Duplicate or split invoice: one PO line billed twice or across several invoices.
- Non-PO charges: freight, surcharges or tax lines with no PO counterpart.
- Vendor or PO reference error: wrong PO number keyed on the invoice.
For each exception, the resolution should be recorded as a coded action (approve with variance, request credit memo, adjust PO, post receipt, reject, hold) plus the free-text note the AP clerk or buyer wrote. That pairing is what lets you train and evaluate an exception-handling agent rather than a classifier. Supplier-facing context on these records is on invoice exception records and invoice reconciliation.
Capture the matching policy as metadata
Two-, three- and four-way matching are company policies, so every record should state which policy applied and with what tolerances. A two-way match compares only PO and invoice; three-way adds the receipt; four-way adds an inspection or quality acceptance step. The same price difference can be "matched" at one company and "blocked" at another.
Ask for a policy table per supplying company or business unit: match type by item category, price tolerance (percent and absolute), quantity tolerance, whether over-receipt is allowed, and whether invoices may post before receipt. The BPI 2019 categories show why: invoice-before-receipt and invoice-after-receipt items follow different paths through the same process [1]. Without this metadata, labels from different suppliers conflict and the model learns noise.
Illustrative record and schema
The record below shows the structure to request, in JSON Lines so each PO line is one UTF-8 object per line [10].
Illustrative example: invented to show structure; it does not describe an available dataset.
{"match_id": "m-000412", "policy": {"match_type": "three_way", "price_tol_pct": 2.0, "price_tol_abs": 50.00, "qty_tol_pct": 0.0, "invoice_before_gr_allowed": false},
"po": {"po_number": "PO-7Q2X", "po_line": 20, "vendor_key": "V-hash-91c3", "item_desc": "Hex bolt M10x40, zinc", "qty": 500, "uom": "EA", "unit_price": 0.42, "currency": "USD", "doc_ref": "po_7Q2X.pdf#p1"},
"receipts": [{"gr_doc": "GR-55810", "gr_line": 1, "qty": 480, "uom": "EA", "posted": "2025-03-04", "doc_ref": "gr_55810.png"}],
"invoice": {"inv_number": "INV-hash-2208", "inv_line": 3, "qty": 500, "uom": "EA", "unit_price": 0.44, "line_total": 220.00, "doc_ref": "inv_2208.pdf#p1", "bbox": [112, 640, 1490, 668]},
"outcome": {"status": "exception", "exception_types": ["quantity_variance", "price_variance"], "price_var_pct": 4.76, "qty_var": 20},
"resolution": {"action": "adjust_po", "resolved_by_role": "ap_specialist", "days_open": 6, "note": "Short shipped 20 EA; vendor price above contract. PO quantity adjusted to 480 to match receiving; invoice paid for actual quantity received."},
"audit": {"change_docs": [{"table": "CDPOS", "field": "MENGE", "old": "500", "new": "480"}]}}
Note what makes it trainable: line-level keys on every document, bounding boxes back to the invoice image, the policy that defines "exception", a coded action, the human note, and the change history.
Masking that preserves the joins
Supplier names, prices and contract terms are confidential, and masking must keep matching keys consistent across all three documents. If the vendor is "V-hash-91c3" on the PO, it must be the same token on the receipt and invoice, and the replaced text in the image must match the replaced value in the structured row. Independent per-document redaction breaks the join and silently destroys the label.
Practical checks to request:
- Deterministic tokenization (keyed hashing) for vendor IDs, PO and invoice numbers, applied identically to images, OCR text and ERP fields.
- Price handling stated explicitly: kept, scaled by a fixed per-vendor factor, or bucketed. Scaling preserves variance percentages; bucketing does not.
- Contact names, emails, phone numbers and bank details on remittance blocks removed in pixels, the OCR layer and PDF metadata. See redacting PII in scanned documents.
- A join-integrity report: share of records where all three documents still resolve to the same PO line after masking.
Evaluating matching and exception agents
Hold out complete transactions, not individual documents, and score matching, extraction and resolution separately. If an invoice is in training and its PO is in test, the split leaks. Split by vendor and by time as well, because templates and tolerance policies change.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Layer | What to score | Typical failure mode |
|---|---|---|
| Extraction | Field and line-item accuracy on PO, receipt, invoice | Line items merged or split across page breaks |
| Linking | Correct PO line for each invoice line | Wrong line chosen when items repeat with different prices |
| Match status | Agreement with posted status under the recorded policy | Correct math, wrong tolerance applied |
| Exception type | Multi-label accuracy per class | Missing-receipt confused with quantity variance |
| Resolution | Action matches human action; note is grounded | Agent overrides a hold that policy required |
Agentic workflows add their own failure points in planning and tool calls [8], so keep resolution trajectories separate from the document set. For field-level ground truth design, see document extraction evaluation sets.
Sourcing checklist for buyers
Brief suppliers with a request that names the documents, the join and the labels, not just a volume. Document the result with a datasheet in the style of Data Cards, covering upstream source, annotation method and intended use [9].
Illustrative example: invented to show structure; it does not describe an available dataset.
- Documents: PO, goods receipt or receiving log, supplier invoice; credit memos where issued.
- Formats: native PDF and scanned images; ERP rows as Parquet or JSON Lines.
- Keys: PO and line, GR and line, invoice and line, masked vendor key.
- Policy metadata: match type, tolerances, invoice-before-receipt rule, per business unit.
- Labels: line-level match status, exception classes, coded resolution, resolution note, days open.
- History: change documents or audit log for PO and invoice edits.
- Mix: share of exceptions versus clean matches; spread of vendors, templates and years.
- Rights: confirmation the supplying company owns the records and that vendor contracts allow use for model training.
Related pages: invoice line-item extraction data, documents paired with system-of-record entries, and table structure recognition data. Broader context is on the document AI data hub, accounting reconciliations and supply chain and logistics datasets.
How SourceX approaches three-way match data
SourceX sources operational datasets, including finance workflow records like these, from US companies on request; nothing is held in stock and a request does not guarantee a match. Buyers describe the data they need, not the businesses, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked. No masking method is perfect. You can describe the linked AP records you need and the process runs Find, Assess, Agree, Transact and Manage.
Sourcing linked PO, receipt and invoice data
SourceX sources operational datasets, including finance and procurement workflow records, from US companies and manages licensing for AI teams wherever they are based. Nothing is contracted until a supplier agrees, and every dataset is rights-reviewed and delivered under a license. Describe your three-way match data needs.
Frequently asked questions
Can I build a three-way match dataset from public data alone?
Only partly. BPI Challenge 2019 gives real matching categories and event sequences [1] [2], and DocILE gives invoice extraction tasks [5], but no public set links real PO, receipt and invoice images at the line level with resolutions.
Is synthetic data enough for exception handling?
Synthetic logs help test pipelines [4], but variance rates, resolution choices and clerk notes are where synthetic data diverges from real AP. See synthetic vs real documents.
Should exception trajectories live in the same dataset?
Keep them linked by match ID but separate. The document set trains extraction and linking; multi-step resolution trajectories with tool calls are an agent dataset with different evaluation needs [8].
Sources
- International Conference on Process Mining (ICPM 2019), "BPI Challenge 2019" (2019). https://icpmconference.org/2019/?p=302
- 4TU.ResearchData, "BPI Challenge 2019 (OCEL)". https://data.4tu.nl/datasets/46a7e15b-10c7-4ab2-988d-ee67d8ea515a
- arXiv (Berti, van der Aalst et al.), "OCEL (Object-Centric Event Log) 2.0 Specification" (2024). https://arxiv.org/pdf/2403.01975
- Zenodo, "Simulated Object-Centric Event Logs (OCEL 2.0) for Order-to-Cash, Procure-to-Pay, Hiring, and Hospital Patient Lifecycle Processes" (2024). https://zenodo.org/records/13879980
- arXiv (Simsa et al.), "DocILE Benchmark for Document Information Localization and Extraction" (2023). https://arxiv.org/pdf/2302.05658
- arXiv, "Invoice Information Extraction: Methods and Performance Evaluation" (2025). https://arxiv.org/pdf/2510.15727
- SAP Help Portal, "Change Documents (CDHDR and CDPOS)". https://help.sap.com/saphelp_nw73/helpdata/en/48/dfd498ab14280de10000000a42189c/content.htm
- arXiv, "SHIELDA: Structured Handling of Exceptions in LLM-Driven Agentic Workflows" (2025). https://arxiv.org/pdf/2508.07935
- Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- jsonlines.org, "JSON Lines". https://jsonlines.org/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.