Skip to content

Document AI data

Bank Statements and Income Documents for Lending Document AI

Quick answer

A useful bank statement dataset for lending document AI pairs real statement images with transaction-level labels (date, description, amount, running balance) and statement-level fields, across many banks, capture conditions and edited copies. Pay stubs need real documents because payroll layouts vary; official W-2 and 1099 layouts are fixed, so synthetic fills cover them well, though payroll-printed substitute copies vary. Source with GLBA reuse limits and IRC 7216 in mind, and redact identifiers in pixels, not only in extracted text.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What a lending IDP team actually needs from statement data

Lending teams need statement data that trains three distinct capabilities: table extraction, cash-flow feature derivation and tamper detection. A parser that reads 95% of rows is still unusable for underwriting if it drops the one NSF fee or misreads a payroll deposit, because downstream features such as average daily balance, deposit regularity and overdraft counts are computed from every row. Research on statement-based credit scoring, such as the AI-BAAM work using six months of transactions from 611 MSME loan applicants, shows why lenders care about complete transaction histories rather than headline balances [4].

Income verification adds a second document family. Pay stubs carry gross and net pay, pay period, pay date, year-to-date totals and deduction lines; W-2s and 1099s carry annual wage, withholding and payer fields in boxes defined by the IRS. A model that cross-checks a stub's YTD gross against deposits on the statement needs both document types from comparable time windows, which is why this page treats them together. Structured transaction feeds are a separate product; see financial transaction data if you need ledger rows rather than documents.

Label schema for statements and pay stubs

The minimum useful label set has a statement header block, a transaction table with row-level fields, and a reconciliation check. Running-balance arithmetic is the strongest free quality signal in this domain: if opening balance plus signed amounts does not equal each row's running balance and the closing balance, either the labels or the source document are wrong. That check catches OCR digit swaps, missed rows at page breaks and, sometimes, edited PDFs.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "doc_id": "stmt_000412",
  "doc_type": "bank_statement",
  "institution_type": "credit_union",
  "capture": "phone_photo",
  "pages": 4,
  "account_holder_block": "[REDACTED_NAME] / [REDACTED_ADDR]",
  "account_number": "[MASKED_LAST4:XXXX]",
  "period_start": "2025-03-01",
  "period_end": "2025-03-31",
  "opening_balance": 2140.18,
  "closing_balance": 1873.52,
  "transactions": [
    {"row": 1, "page": 1, "bbox": [72, 310, 540, 324], "date": "2025-03-03",
     "description_raw": "ACH CREDIT PAYROLL [EMPLOYER_TOKEN] PPD",
     "amount": 1612.40, "direction": "credit", "running_balance": 3752.58,
     "category": "payroll_deposit"},
    {"row": 2, "page": 1, "bbox": [72, 326, 540, 340], "date": "2025-03-04",
     "description_raw": "POS DEBIT 0303 [MERCHANT_TOKEN]",
     "amount": -48.17, "direction": "debit", "running_balance": 3704.41,
     "category": "card_purchase"}
  ],
  "reconciles": true,
  "tamper_label": "none",
  "redaction_method": "pixel_blackout+token_replace"
}

Keep description_raw as printed, noise included, and put normalized merchant or counterparty values in a separate field. Bounding boxes per row let you evaluate table structure recognition, not just field values; the table structure recognition data page covers cell and row annotation conventions. For pay stubs, label each earnings and deduction line with current and YTD amounts, because YTD-to-current inconsistencies are both an extraction error signal and a fraud signal. The broader key-value extraction labels page explains entity linking for header fields.

Coverage: the formats that break statement parsers

Format coverage matters more than raw volume, because statement parsers fail on layout variety, not on repetition of a layout they already handle. A set of 50,000 statements from three large banks will train a model that collapses on a regional credit union's two-column layout. Write coverage targets into the request by institution type, layout family and capture condition.

Illustrative example: invented to show structure; it does not describe an available dataset.

Coverage dimensionWhy it mattersExample target in a request
Institution mixLayouts differ by core banking vendor and print templateNational banks, regional banks, credit unions, neobanks
Statement lengthRows split across page breaks, carried-forward balances1-page through 12+ page statements
Account typeBusiness statements carry wire, ACH batch and merchant settlement rowsConsumer checking, savings, business operating accounts
CapturePhone photos skew and shadow tables; faxes lose decimalsNative PDF, scanned, photographed, re-printed
OriginOnline-banking exports differ from mailed statementsPortal PDF, mailed statement scan, aggregator render
IntegrityEdited PDFs are the case your fraud model must seeLabeled unmodified vs altered, with altered fields marked
Income docsPayroll providers render stubs differentlySeveral payroll systems, hourly and salaried, multi-state withholding

Capture degradation is its own topic; degraded document images covers fax, photocopy and phone-photo artifacts. Altered statements, including font substitution and recomputed balances, belong in a dedicated set described on document tampering and fraud detection data. Mortgage teams classifying whole loan files should also see mortgage loan file document data.

Where synthetic statements are enough and where they fail

Synthetic data covers fixed-layout tax forms and early pipeline work well, and fails on real transaction-description noise and bank-specific quirks. Open synthetic sets exist: the AgamiAI Indian Bank Statements set on Hugging Face is synthetic, Apache-2.0, ships PDF and JSON, and its card notes that real business banking documents are scarce due to privacy while listing fraud, AML and credit decisioning as out of scope [3]. SynFinTabs generated about 100,000 synthetic financial tables, and its authors still built a small real-world test set to check transfer [5].

Generated statements tend to use clean merchant strings, consistent date formats and well-behaved page breaks. Real descriptions carry truncated merchant names, ACH company IDs, card-terminal codes and bank-specific abbreviations, which is exactly what categorization and payroll-detection models must learn. The official W-2 and 1099 layouts are set by the IRS, so synthetic filled forms cover them well, but employee copies are often substitute statements printed by payroll providers in their own formats, so include some real ones; pay stubs vary by payroll provider and employer configuration, so that is where real documents earn their cost. A practical split is synthetic data for W-2s and pretraining, real documents for statements, pay stubs and the evaluation set; synthetic vs real documents covers the broader trade-offs.

Financial privacy: GLBA, IRC 7216 and redaction in pixels

Bank statements and tax documents carry legal restrictions that follow the data, so the origin of each document decides what a licensor can do with it. Under the Gramm-Leach-Bliley Act privacy rule, a business that receives nonpublic personal information from a financial institution generally steps into the shoes of that institution and faces limits on reusing and redisclosing it [1]. Regulation P section 1016.11 sets those redisclosure and reuse limits [2]. Statements a lender collects from consumer applicants are its customers' nonpublic personal information, so sharing them with a nonaffiliated licensee raises notice and opt-out questions; a business's own operating-account statements fall outside GLBA's consumer scope and are often the cleaner starting point.

Tax return information held by a tax preparer is a separate case: Internal Revenue Code section 7216 and its regulations restrict preparers from using or disclosing return information for purposes other than preparing returns. Treat W-2s, 1099s or 1040s offered by a preparer or tax software channel as a counsel question, not a redaction question, because de-identification alone may not settle it. Check the consent and exception rules before accepting tax documents from that channel.

Redaction must cover the image, not just the OCR layer. A statement with names blanked in the JSON but still visible in the page image, the PDF text layer or the XMP metadata is not de-identified. NIST SP 800-188 describes de-identification techniques and cautions that traditional methods have inherent limits [8], so treat masked account numbers, tokenized employer names and blacked-out addresses as risk reduction, and keep the method on record. The redacting PII in scanned documents guide covers pixel, text-layer and metadata redaction in detail.

Building an evaluation set for statement parsing

A statement-parsing evaluation set should score row recall, field accuracy and reconciliation, not just character error rate. Public form benchmarks help with methodology but not with this document type: FUNSD has 199 noisy scanned forms labeled for entity linking [6], and RealKIE covers enterprise documents such as SEC filings and invoices [7], but neither contains transaction tables.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Row-level recall and precision: missed and phantom transactions per statement, reported separately for page-break rows.
  • Field exact match: date, amount and sign per row; normalized string match for descriptions.
  • Reconciliation pass rate: share of statements where extracted rows reproduce every running balance and the closing balance.
  • Header accuracy: period dates, opening and closing balance, masked account identifier.
  • Stratified slices: results by institution type, capture condition and page count, so a regression on photographed credit-union statements is visible.
  • Income cross-check: agreement between pay-stub net pay and matching statement deposits within a stated tolerance.

Hold the evaluation set out from any vendor that also supplies training data, and freeze it per model version. The document extraction evaluation ground truth page covers adjudication and versioning.

How SourceX approaches statement and income document requests

SourceX sources operational datasets from US companies on request, and documents and finance workflows are among the kinds of data it looks for; nothing is held in stock, and a request does not guarantee a match. Buyers describe the data, such as statement types, layouts, capture conditions and label fields, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery.

Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows after an executed agreement and supplier approval. Teams in fintech software and mortgage operations can start from the buyer page with the coverage table above as their request outline, and the Document AI data hub lists related document types.

Request bank statement and income document data

If your lending or IDP team needs real statements or pay stubs with transaction- and field-level labels, describe the institutions, layouts, capture conditions and label schema you need. SourceX runs the process from Find and Assess through Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Start your request at sourcex.si/buyers.

Sources

  1. Federal Trade Commission, "How To Comply with the Privacy of Consumer Financial Information Rule of the Gramm-Leach-Bliley Act". https://www.ftc.gov/business-guidance/resources/how-comply-privacy-consumer-financial-information-rule-gramm-leach-bliley-act
  2. Consumer Financial Protection Bureau, "12 CFR 1016.11 Limits on redisclosure and reuse of information". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
  3. Hugging Face, "AgamiAI Indian Bank Statements dataset card". https://huggingface.co/datasets/AgamiAI/Indian-Bank-Statements/blob/main/README.md
  4. arXiv, "AI-BAAM: AI-Driven Bank Statement Analytics as Alternative Data for Malaysian MSME Credit Scoring" (2025). https://arxiv.org/pdf/2510.16066
  5. arXiv, "SynFinTabs: A Dataset of Synthetic Financial Tables for Information and Table Extraction" (2024). https://arxiv.org/pdf/2412.04262
  6. arXiv (Jaume, Ekenel, Thiran), "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
  7. arXiv, "RealKIE: Five Novel Datasets for Enterprise Key Information Extraction" (2024). https://arxiv.org/pdf/2403.20101
  8. National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data