Skip to content

Document AI data

Document Tampering and Fraud Detection Data: Altered Invoices, Statements and Records

Quick answer

A usable document fraud detection dataset pairs documents confirmed as manipulated by an investigation with genuine controls from the same issuers and time period. Each tampered file needs labels for the outcome, the altered fields or regions, and the method. The original file must also be kept intact, including PDF metadata, fonts and incremental saves. Public sets mostly cover synthetic ID documents and small receipt collections, so altered invoices, bank statements and business records usually have to come from licensed operational sources.

By SourceX Editorial · Updated

What public tampering datasets cover, and where they stop

Public forgery benchmarks are useful for prototyping, but they cluster around identity documents and receipts; others cover text tampering in documents at scale [7]. not the invoices and statements that lenders, insurers and AP teams actually review. IDNet is one of the largest examples. Its authors released synthetic identity documents because MIDV-500 and MIDV-2020 lacked enough fraud variety [1]. The Zenodo record lists about 837K images across 20 document types [2], and every one of them is generated, not collected from a real fraud case.

Receipt forgery has two academic sets. The ICDAR 2023 "Find it again" dataset takes 988 SROIE receipts and annotates 163 with realistic fraudulent edits [3]. A La Rochelle benchmark provides 1,969 receipt images with OCR output for fraud detection [4]. Both work well for pixel-level forensics research, but they are small, single-domain and mostly image-based. That means they cannot teach a model about native-PDF evidence such as object streams, font subsets or edit history.

The gap is commercial documents. These include vendor invoices with a changed remit-to account, bank statements with inflated deposits, pay stubs with edited gross pay, repair estimates with added line items, and certificates of insurance with changed dates. Before you build on any public set, check its license terms on our guide to public document AI datasets and commercial use.

Labels that make a tampered document trainable

The label that matters is an investigation outcome tied to specific alterations, not the extracted content fields. That is the difference between this data and the extraction sets covered in the Document AI data hub. A "suspected" flag from a rules engine is not ground truth. Ask whether each positive came from a closed SIU, AP-audit or underwriting investigation, and record who confirmed it.

Useful label layers for each document:

  • Outcome: confirmed fraud, confirmed genuine, unresolved, or excluded. Keep unresolved cases separate rather than dropping them, because they show where your model will be uncertain in production.
  • Altered region: page number plus a bounding box or polygon for each changed area, in PDF user-space coordinates for native PDFs and pixel coordinates for images.
  • Altered field: which business field changed (amount, date, payee, IBAN or routing number, employer name, line item), with the original value where the investigation established it.
  • Method: edited native PDF (text object replaced), pasted image patch, re-typed or fabricated template, mixed pages from different genuine documents, print-scan laundering, or screenshot of an edited web statement.
  • Evidence basis: how the alteration was proven, for example issuer confirmation, bank verification, metadata inconsistency, or a font mismatch found on review.

Label noise is the biggest quiet risk here. Benchmarks that are widely trusted still carry label error rates of at least 3.3% on average, enough to change model rankings [6]. Fraud labels are noisier still, because an unproven case often gets closed as "genuine." Require a written definition of each outcome value and a double-review rate for a sample of positives.

Keep file-level evidence intact

Many document tamper detectors use evidence that disappears the moment a file is rasterized, re-saved or passed through a redaction tool. For native PDFs, that evidence includes the Info dictionary and XMP metadata (Producer, Creator, CreationDate, ModDate), incremental-update sections and multiple %%EOF markers, cross-reference tables, embedded and subset fonts, and text objects whose glyphs don't match surrounding font metrics. For images, it includes EXIF, JPEG quantization tables and double-compression traces that error-level and splicing detectors rely on.

Specify in your request that originals are delivered byte-for-byte alongside any de-identified rendering, with SHA-256 hashes for each file. Then check what de-identification does to the evidence. Redacting a payee name by drawing a box over it, then flattening the file, can erase the incremental saves you wanted to learn from. See redacting PII in scanned documents for how pixel, OCR-layer and metadata redaction interact. Agree in advance which metadata fields are kept, generalized or removed.

Build a matched genuine control set

Class imbalance in confirmed fraud is extreme. In one Medicare study, the fraud rate in enforcement-labeled data was roughly 0.038% to 0.074% [5]. Document fraud in AP or lending files also tends to be rare, so a dataset made only of positives produces a detector that learns issuer or template identity instead of tampering.

Match controls to positives on issuer or template family, document type, date range, capture channel (native PDF upload, phone photo, scan) and producing software. If the tampered bank statements come mostly from three issuers' PDF generators, the genuine statements should come from the same generators and the same months. For statement-specific layouts and fields, see bank statements and income documents for lending document AI. For capture artifacts that can be mistaken for tampering, see degraded scans, faxes and phone photos.

Write down the ratio you plan to train on, and evaluate separately at the production base rate. A model tuned on a 1:1 split can see its precision collapse at 1:1,000.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueNotes
doc_idINV-000481Stable pseudonymous ID
doc_typevendor_invoiceFrom your taxonomy
file_sha2569f2c...e1Hash of the untouched original
capture_channelnative_pdf_uploadnative_pdf, scan, phone_photo, email_body
producer_softwaregeneralized: "office suite"Kept, generalized or removed per agreement
outcomeconfirmed_fraudconfirmed_fraud, confirmed_genuine, unresolved
confirmed_byAP audit, issuer callbackEvidence basis
altered_fieldsremit_bank_account; total_amountBusiness fields changed
altered_regionsp1:[412,188,560,204]; p1:[470,612,560,628]PDF user-space boxes
methodedited_native_pdf_incremental_saveMethod taxonomy value
control_group_idCG-07Links to matched genuine documents
investigation_closed2025-Q3Quarter only

Synthetic tampering: where it helps and where it misleads

Synthetic tampering helps with augmentation and localization pretraining, but it does not reproduce the techniques real fraudsters use. Scripts that paste random patches or swap digits teach a model to find the generator's artifacts. Real cases often involve template fabrication, careful font matching, print-scan laundering, or edits made in the same software that produced the genuine file.

A practical split is to use synthetic edits on genuine documents to pretrain localization, then fine-tune and, above all, evaluate on investigation-confirmed cases. Keep synthetic and real positives clearly flagged so evaluation never mixes them. The tradeoffs are covered in more detail in synthetic vs real documents for document AI, and evaluation design is covered in outcome-labeled evaluation data.

Rights, privacy and confidentiality checks for fraud records

Fraud records need more legal review than ordinary business documents, so confirm their status before any licensing discussion. Files tied to open claims, litigation or referrals may be under legal hold or subject to SIU confidentiality. The fraudster's identity, the victim's account numbers and the investigator's notes all need separate handling.

Questions buyers should put to any supplier:

  1. Is each case closed, and is any file under a legal hold or a law-enforcement referral that limits disclosure?
  2. Who owns the document: the submitting party, the issuer it imitates, or the company that received it?
  3. How are account numbers, names and addresses removed or replaced, and does that method preserve the altered regions and file metadata?
  4. Does the dataset include health records (for example, claim attachments)? If so, HIPAA de-identification applies.
  5. Which uses does the license allow: training, evaluation, or both? And does it allow derived synthetic data?

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

How SourceX approaches tampered-document requests

SourceX sources operational datasets from US companies and manages the commercial process, from licensing agreements to ongoing purchases. Data is sourced on request, not held in stock, and a request does not guarantee a match. You describe the data, not the businesses. SourceX looks for US businesses that hold it, and every release is approved by the supplying company.

Each dataset is reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Names, emails, phone numbers and account numbers are removed or replaced before delivery. The method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Teams scoping a fraud set can start with the SourceX buyer request, and insurers can also see buyers by industry: insurance. For general request structure, use the document dataset requirements spec.

Request document fraud detection data

Describe the document types, the alteration methods you need covered, the label layers, and the genuine control set you want matched. SourceX then works through Find, Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Submit your document fraud detection data request.

Sources

  1. arXiv, "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection" (2024). https://www.arxiv.org/pdf/2408.01690
  2. Zenodo, "IDNet dataset record" (2024). https://zenodo.org/records/13852207
  3. L3i, La Rochelle University, "Find it again: a receipt forgery dataset (ICDAR 2023)" (2023). https://l3i.univ-larochelle.fr/app/uploads/sites/12/2024/05/ICDAR_2023_Find_it_again.pdf
  4. HAL, La Rochelle University, "Receipt dataset for fraud detection" (2019). https://hal-univ-rochelle.archives-ouvertes.fr/hal-02316349v1
  5. medRxiv, "Utilization Analysis and Fraud Detection in Medicare via Machine Learning" (2024). https://www.medrxiv.org/content/10.1101/2024.12.30.24319784.full.pdf
  6. arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  7. CVPR 2023, "Towards Robust Tampered Text Detection in Document Image: New Dataset and New Solution". https://arxiv.org/pdf/2303.00310

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data