Document AI data
Receipt Line-Item Data for Expense AI: Thermal, Crumpled and Photographed Receipts
Quick answer
A receipt dataset with line items pairs each receipt image with structured labels for merchant, date, subtotal, tax, tip, total, currency and every purchased line (description, quantity, unit price, amount). Public sets such as CORD and SROIE are useful baselines but are small, regional and not labeled by capture condition. Expense-AI teams usually need real employee-submitted photos, including faded thermal paper and folded receipts, linked to the expense report that approved or flagged them.
By SourceX Editorial · Updated
What public receipt datasets cover, and where they stop
Public receipt sets give you a schema and a benchmark, not production coverage. CORD from NAVER CLOVA pairs Indonesian receipt photos with a parse schema that includes menu line items (name, count, price) and is released under CC-BY-4.0 [1]. It is small for production work, and listed counts and splits differ between the GitHub repository and third-party zoos, so pin a version before you report scores [1][4].
SROIE, the ICDAR 2019 scanned-receipt set, is the other common baseline; its key-information task targets header fields such as company, date and total rather than line items, so confirm the release you use and do not expect it to train a line-item parser. CORU adds regional diversity with 20,000+ annotated receipts from Egyptian retail [3], and a 2020 study built its own 65-receipt, 711 line-item set because it found no public line-item receipt dataset at the time [2].
For a US expense product, that leaves gaps in currency, tax regime, merchant mix (fuel, hotel folios, rideshare, parking) and capture quality. See which public document AI datasets allow commercial use before you build on any of them, and check each card's license field yourself: an audit of popular hosting sites found license omission above 70% and license errors above 50% [6][7].
Capture conditions that break receipt extraction in production
Real receipts fail in ways clean scans never show, so a training or eval set should label capture condition per image. The failure modes worth stratifying are:
- Thermal fade: low-contrast or partially erased text on older thermal paper, often worst on totals printed at the bottom.
- Folds and crumpling: creases that split a line-item row across two text lines or bend the price column.
- Long receipts: grocery or hardware receipts photographed in two or three overlapping shots, or cropped mid-list.
- Multiple receipts per photo: a hotel folio and a parking stub on one desk shot.
- Lighting and angle: shadows from the phone, glare on glossy paper, perspective skew.
- Handwritten tips and totals: restaurant slips where the authoritative total is written by hand.
The degraded document capture conditions guide covers the general taxonomy; for receipts, the practical point is that synthetic augmentation (blur, noise) does not reproduce thermal fade or tip handwriting well. The synthetic vs real documents comparison explains where generated data breaks.
Expense-report outcomes as labels
The most valuable labels for expense AI come from the expense workflow, not from annotators. When a receipt image stays linked to its expense line, you inherit the submitted amount, the approved amount, the expense category, the cost center, the policy flags raised and the approver decision. Those fields support categorization, duplicate detection and policy-compliance agents that a pure OCR set cannot train.
Approval history behaves like a process log. The public BPI Challenge 2020 declarations data shows how travel expense claims move through approval steps as recorded events [5]; a commercial equivalent attaches those events to the receipt images themselves. Treat system-of-record values with care, because approved amounts can legitimately differ from printed totals (partial reimbursement, per-diem caps, personal items removed). The documents paired with system-of-record entries guide covers that label-noise problem, and spend classification training data covers mapping line descriptions to category taxonomies.
An illustrative receipt line-item record
A usable record keeps the image, the printed values, the line items and the expense outcome in separate, typed blocks.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"receipt_id": "r_000184",
"image_files": ["r_000184_p1.jpg", "r_000184_p2.jpg"],
"capture": {"source": "mobile_photo", "paper": "thermal", "fade": "moderate", "folded": true, "pages": 2},
"merchant": {"name": "Hardware store (retained)", "category_hint": "supplies", "country": "US", "state": "OH"},
"transaction": {"date": "2025-03-14", "time": "16:42", "currency": "USD", "payment": "card_last4_masked"},
"line_items": [
{"line": 1, "desc": "DECK SCREW 2.5IN 1LB", "qty": 2, "unit_price": 9.48, "amount": 18.96, "bbox": [112, 420, 880, 452]},
{"line": 2, "desc": "PAINTER TAPE 1.88", "qty": 1, "unit_price": 6.97, "amount": 6.97, "bbox": [112, 458, 880, 490]}
],
"totals": {"subtotal": 25.93, "tax": 1.88, "tip": null, "total": 27.81},
"expense_outcome": {"category": "Office/Facilities supplies", "submitted": 27.81, "approved": 27.81, "policy_flags": [], "decision": "approved"},
"redaction": {"card_digits": "masked", "loyalty_id": "removed", "employee_name": "replaced", "method_ref": "deid_v2"}
}
Check that line-item amounts reconcile to the subtotal and that subtotal plus tax plus tip equals the total; records that do not reconcile should carry a reason code (discount line, unreadable row, multi-page gap) rather than silent edits. The same reconciliation logic appears in invoice line-item extraction data and key-value extraction labels.
Scoring receipt models field by field
A single document-level accuracy number hides the errors that cost money, so score each field family separately. Use exact match after normalization for total, tax and date; normalized string similarity for merchant name; and line-level precision and recall for line items, matching on description plus amount. Report results by capture condition, merchant category and receipt length, because a model that reads short restaurant slips well can still miss half the rows on a two-photo grocery receipt.
Hold out merchants, not just receipts, to avoid template leakage: chains print identical layouts, and a random split rewards memorizing them. Keep a fixed eval set separate from any training purchase, with its version and redaction method documented.
Privacy and redaction on receipts
Receipts carry fewer personal fields than invoices, but the ones they carry are sensitive. Mask card numbers beyond what the receipt already truncates, remove loyalty and membership IDs, phone numbers and email addresses printed for e-receipts, and replace employee and cardholder names. Merchant names, store addresses and item descriptions usually stay, because removing them destroys the extraction task.
Watch for edge cases: pharmacy receipts can reveal health information, hotel folios list guest names and room numbers, and rideshare receipts include pickup and drop-off addresses. Decide in the license whether those subtypes are excluded or redacted further.
Buyer checklist for a receipt line-item request
Describe the data precisely so a supplier can tell whether they hold it.
- Volume range and receipt mix by merchant category (meals, lodging, fuel, ground transport, supplies).
- Capture sources: mobile photo, email e-receipt, scanned paper; target share of thermal and folded receipts.
- Label depth: header fields only, or full line items with bounding boxes.
- Linked expense outcomes: category, approved amount, policy flags, approver decision.
- Currencies, countries and tax regimes needed.
- Date range and whether ongoing deliveries are wanted.
- Redaction requirements and the evidence you expect (method description, sample check).
- Allowed uses: training, fine-tuning, evaluation, or all three.
How SourceX handles receipt line-item requests
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Nothing is held in stock and a request does not guarantee a match: you describe the receipts and labels you need, SourceX looks for US businesses that hold them, and every release is approved by the supplying company. You can start a request on the SourceX buyer page.
Each dataset is rights-reviewed for ownership and consents, with diligence materials on source, rights, preparation and allowed use prepared per dataset. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. For the record type in general, see licensing invoices and receipts for AI training and what makes invoices and receipts valuable for AI; the Document AI data hub and the AI data hub list related document types.
Request receipt line-item data for expense AI
Describe the receipt mix, capture conditions, line-item labels and expense outcomes you need, and SourceX will look for US companies that hold that data. Pricing and allowed uses are agreed in a license, delivery runs through private, access-controlled workflows after an executed agreement and supplier approval, and nothing is contracted until a supplier agrees. Start a buyer request.
Sources
- NAVER CLOVA AI, "CORD: A Consolidated Receipt Dataset for Post-OCR Parsing (GitHub repository)". https://github.com/clovaai/cord
- arXiv, "Understanding Scanned Receipts" (2020). https://arxiv.org/pdf/2005.01828
- Hugging Face (abdoelsayed), "CORU dataset card (README.md)". https://huggingface.co/datasets/abdoelsayed/CORU/blob/f94dff8081f54ef05ba7586a3db2b10ed17942d3/README.md
- Voxel51, "Consolidated Receipt Dataset (dataset zoo entry)". https://docs.voxel51.com/dataset_zoo/datasets_hf/consolidated_receipt_dataset.html
- 4TU.ResearchData, "BPI Challenge 2020: International Declarations" (2020). https://data.4tu.nl/articles/_/12687374/1
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.