Skip to content

Industry-specific operational data

Mortgage loan file document data for classification and extraction models

Quick answer

Mortgage document AI needs real loan packages, not generic forms: merged PDFs where each page carries a document type (URLA, Loan Estimate, Closing Disclosure, appraisal, paystubs, title, note), document boundaries and version markers are labeled, and key fields have ground truth tied to the loan's system of record. Buy for spread across lenders, channels, loan programs and vintages, and require GLBA-grade redaction of text and images plus a documented licensing basis from the originating lender.

By SourceX Editorial · Updated

What a mortgage loan file actually contains

A closed-loan file typically holds dozens of document types, often running to several hundred pages, assembled across origination, processing, underwriting and closing, usually exported as one or a few merged PDFs from a loan origination system. The anchor is the Uniform Residential Loan Application (Form 1003), which Fannie Mae and Freddie Mac redesigned and paired with the Uniform Loan Application Dataset (ULAD), a mapping of its fields to the MISMO v3.4 reference model. That mapping matters for labels: a URLA field such as gross monthly income has a MISMO data point you can use as a canonical key. Confirm current form versions and effective dates on the GSE URLA pages before fixing your taxonomy.

Disclosures follow the CFPB's TILA-RESPA Integrated Disclosure (TRID) rule. Regulation Z at 12 CFR 1026.37 and 1026.38 defines Loan Estimate and Closing Disclosure content, and 1026.19 governs timing and revisions, which is why one file often holds several LE and CD versions. The CFPB's guide to the LE and CD forms is the practical reference for how each page is laid out.

Around those sit income and asset evidence (W-2s, paystubs, 1040s, 4506-C transcript requests, bank statements), the appraisal (commonly the Uniform Residential Appraisal Report, Form 1004), title commitment, flood certification, hazard insurance declarations, the note, the security instrument, and conditions correspondence. Income and bank-statement pages have their own owner guide on bank statements and income documents for lending document AI; this page covers the package as a whole.

Why public datasets do not cover mortgage packages

Public form datasets are too small and too generic for mortgage IDP. FUNSD, a common form-understanding benchmark, has 199 annotated scanned forms from domains such as marketing and scientific reports [3]. RealKIE's five enterprise extraction sets cover SEC S-1 filings, NDAs, UK charity reports, FCC invoices and resource contracts, with no URLA, TRID or appraisal pages [4]. Public HMDA releases are structured loan-level records, not document images, so they cannot train a page classifier.

The failure modes that hurt production models only appear in real files: fax-degraded paystubs, rotated pages, borrower-signed URLA addenda interleaved with lender-generated pages, a revised CD stapled after the original, and appraisals with embedded photos and sketches. Synthetic packages rarely reproduce the investor-specific stacking order a post-closing team actually uses.

Labels to specify before you request data

Specify labels at three levels: page, document and field. Most mortgage IDP pipelines split a merged PDF into documents, classify each, then extract fields, so each stage needs its own ground truth. General techniques are covered on document classification training data and page-stream segmentation data; the mortgage-specific part is the taxonomy.

  • Page level: document type from a fixed taxonomy, page index within the document, orientation, blank/separator flag, and OCR confidence.
  • Document level: start and end page in the merged PDF, version (LE v1, revised LE, initial CD, final CD), signed versus unsigned, and the stacking position used by the lender or investor.
  • Field level: key values with bounding boxes and a normalized value (for example, loan amount, note rate, cash to close, appraised value, borrower base income), ideally reconciled to the LOS record. Key-value extraction labels explains annotation conventions; pairing values with system records is covered in documents paired with system-of-record entries.

Stacking orders differ by investor and lender, so ask suppliers to deliver their native document taxonomy plus a mapping table to yours rather than relabeling everything to one scheme.

Illustrative example: invented to show structure; it does not describe an available dataset.

{"loan_ref": "L-000184", "file": "closed_package_01.pdf", "page": 42,
 "doc_type": "CLOSING_DISCLOSURE", "doc_version": "final_cd", "doc_start_page": 40, "doc_end_page": 44,
 "stack_position": "investor_order_17", "channel": "wholesale", "program": "FHA", "vintage": "2023",
 "fields": [{"name": "cash_to_close", "value": "[REDACTED_AMOUNT_BAND_3]", "bbox": [412, 588, 530, 604],
             "source_of_truth": "los_closing_record"}],
 "redaction": {"text": "token_replace", "image": "black_box", "signatures": "masked"}}

Coverage and spread that drive robustness

Spread matters more than raw page count. A model trained on one lender's retail conventional loans from one LOS template will misread wholesale broker packages, correspondent files with another lender's disclosures, and FHA or VA files with program-specific forms. Request a coverage matrix up front.

Illustrative example: invented to show structure; it does not describe an available dataset.

DimensionWhat to ask forWhy it matters
ChannelRetail, wholesale, correspondentDifferent document generators and stacking habits
ProgramConventional, FHA, VA, jumboProgram-specific forms and conditions
VintagePrior and redesigned URLAOld and redesigned Form 1003 layouts differ
Disclosure versionsFiles with revised LE and multiple CDsVersion disambiguation is a common extraction error
Image qualityNative PDF, scanned, faxedOCR robustness
Appraisal formsCurrent and prior appraisal report versionsAs of October 2026 the GSEs are moving to a redesigned URAR; check their publications for timing before fixing your target

For condition-level reasoning on the same files, see mortgage underwriting conditions and condition-clearing records.

Privacy and licensing basis for borrower files

Every loan file is dense with nonpublic personal information: SSNs, account and loan numbers, income, signatures, and property addresses. Under Regulation P, a recipient of NPI from a nonaffiliated financial institution can use and redisclose it only within limits tied to how it was received, so the lender's basis for licensing the data has to be documented [1][2]. Ask for redaction of both the OCR text layer and the page image, since a clean text layer over an unredacted image still leaks.

Practical checks: confirm that signatures, handwritten notes and appraisal photos of occupied homes are handled; confirm whether amounts are kept, banded or perturbed, because extraction ground truth depends on that choice; and get the de-identification method written down. The process is walked through in how to de-identify underwriting files. This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Buyer checklist for a mortgage document data request

Describe the data you need, not the lenders you want. A strong request fits on one page.

Illustrative example: invented to show structure; it does not describe an available dataset.

  1. Task: splitting, classification, field extraction, evaluation, or condition-matching agents.
  2. Document taxonomy and the minimum set of required types (URLA, LE, CD, 1004, note, title).
  3. Label levels needed (page, document, field) and whether field values must reconcile to LOS records.
  4. Coverage matrix: channel, program, vintage and image-quality targets.
  5. Redaction expectations for text, image and signatures, and how amounts are treated.
  6. Allowed uses (training, evaluation, internal benchmarking) and delivery format, such as PDFs plus JSON Lines labels.
  7. Documentation: a datasheet or Data Card covering source, annotation method and known gaps [5].

Teams building evaluation sets should also read document extraction evaluation ground truth. Broader industry context is on the industry-specific operational data hub and the mortgage operations buyer page.

How SourceX sources mortgage loan file data

SourceX sources operational datasets from US companies, including document and finance workflows, on request rather than from stock, and a request does not guarantee a match. It looks for US businesses that hold the described data, every release is approved by the supplying company, and each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details such as names, account numbers and phone numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe your mortgage document data needs to SourceX; generic underwriting files are covered on license underwriting files for AI training, and RAG uses on document AI and RAG datasets.

Request mortgage loan file document data

SourceX runs the process as Find, Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval, with diligence materials prepared per dataset. Start a mortgage document data request.

Frequently asked questions

Can HMDA data replace loan documents for training?

No. Public HMDA data is structured loan-level reporting, so it can support tabular modeling but has no page images or layouts for classification or extraction.

Should redacted values still be usable as extraction labels?

Only if the redaction method preserves them. Token replacement or banding can keep field position and type, which is enough for localization and classification but not for exact-value accuracy metrics; agree on the method before labeling.

Do I need both old and redesigned URLA versions?

If your model will process seasoned loans, servicing transfers or post-closing audits, yes. Older files can carry the prior Form 1003 layout, so a model trained only on the redesigned form will see out-of-distribution applications.

Sources

  1. Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
  2. Federal Trade Commission, "How To Comply with the Privacy of Consumer Financial Information Rule of the Gramm-Leach-Bliley Act". https://www.ftc.gov/business-guidance/resources/how-comply-privacy-consumer-financial-information-rule-gramm-leach-bliley-act
  3. Jaume, Ekenel, Thiran, "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
  4. arXiv, "RealKIE: Five Novel Datasets for Enterprise Key Information Extraction" (2024). https://arxiv.org/pdf/2403.20101
  5. Pushkarna, Zaldivar, Kjartansson, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data