Industry-specific operational data
CMS-1500 and UB-04 claim form images for claims intake OCR
Quick answer
A usable CMS-1500 or UB-04 dataset for OCR is a set of real scanned paper claims paired with the keyed claim record each image produced, aligned at the box or form-locator level. Specify the form version (CMS-1500 02/12, UB-04/CMS-1450), the print-quality mix (red dropout originals, black-and-white copies, faxes), the handwriting share, how ground truth was keyed and checked, and how identifiers were removed from the pixels themselves. Public form benchmarks do not cover these forms.
By SourceX Editorial · Updated
This guide is for document AI leads at payers, TPAs and mailroom vendors who are training or evaluating extraction models. It sits in the industry-specific operational data hub and pairs with our general guide to OCR ground truth aligned to real business scans.
Why claim forms need their own dataset spec
Claim forms need their own spec because their schema is fixed by Medicare manuals and a standards body, and generic form corpora do not reproduce it. CMS-1500 completion instructions and print specifications sit in Chapter 26 of the Medicare Claims Processing Manual (Pub. 100-04), with item-level guidance for Items 1 through 33 of the 02/12 version. For the institutional UB-04 (CMS-1450), the reference is Chapter 25, and the National Uniform Billing Committee (NUBC) owns form design and the approved code lists.
Chapter 25 groups its instructions by form locator ranges that run through FL 81, and its section numbers have shifted between revisions. Pin the revision you label against. If you want a supplier to find this material for you, start with a precise brief; SourceX works with AI data buyers on requests like this. The most widely cited public form benchmark, FUNSD, contains 199 noisy scanned forms from domains such as marketing and scientific reports [4]. It is useful for pretraining key-value linking, but it has no service-line grids, no diagnosis pointers and no dropout ink.
As of October 2026, confirm the current CMS-1500 and UB-04 versions with CMS or your MAC before you write the spec. Forms printed under an older version can still arrive in a mailroom, so ask suppliers to tag version per image rather than assume one.
Dropout ink, copies and faxes: the print-quality mix to request
The print-quality mix matters more than raw image count because dropout behavior decides what your preprocessing sees. Medicare's print specification calls for forms printed in Flint OCR Red J6983 or an exact match, so OCR scanners can drop the form lines and read only the entered data. Downloaded and printed copies may not reproduce the scale and OCR color, which is why they are not accepted for Medicare submission.
Medicare contractor guidance tells providers to fill forms in black ink only, because scanners drop red, and to avoid worn toner, faded ribbons and dot matrix output, which breaks characters. It asks for uppercase text and warns that misaligned data may not be read, causing denials or incorrect payment. Every one of those instructions describes a failure mode your model will meet in real mail: red text entries that vanish, black-and-white printed forms where the grid survives dropout, and data shifted across box boundaries.
Ask suppliers to label each image with these strata rather than reporting a single quality score:
- Form stock: red dropout original, black-and-white print, photocopy, fax, or downloaded PDF printout.
- Entry method: practice-management printout, typewriter, handwriting, or mixed (typed claim with handwritten corrections).
- Capture path: mailroom production scanner (resolution, bit depth, color or bitonal), fax server TIFF, or rescanned archive.
- Registration: in-box, shifted by a fraction of a line, or shifted by a full line or column.
- Attachments: whether the envelope also held an EOB, medical records or a cover letter, and how pages were split.
Paper claims that still reach a payer are not a random sample of claims. Medicare accepts paper claims mainly from providers that qualify for a waiver from the Administrative Simplification Compliance Act (ASCA) electronic submission requirement. Ask suppliers where their paper volume comes from, such as small practices, out-of-network billers or non-Medicare lines, and set coverage targets so your test set reflects your own mail rather than a supplier's convenience.
Box-level ground truth from keyed adjudication data
Ground truth should come from the claim record that was keyed from each image and then adjudicated, aligned back to boxes, with the keying error rate measured. A mailroom vendor or payer usually holds three layers: the image, the data-entry record (often converted to an 837P or 837I for the claims system), and the adjudicated claim with edits applied. The keyed record is closest to what was on paper; the adjudicated record may contain corrections that never appeared on the form.
Ask which layer each label came from and whether double-key verification or a QA sample was used. If a supplier only has adjudicated data, expect label noise where examiners corrected NPIs, dates or units. For a deeper treatment of matched claim outcomes, see claim denial prediction data built from 837 and 835 pairs.
Geometry matters as much as values. Request bounding boxes per field and per service line, in a documented format such as JSON with pixel coordinates, or a layout format such as ALTO or PAGE XML if your pipeline already consumes one. Without coordinates, you can train end-to-end extraction but cannot measure whether errors come from detection, recognition or line association.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field (CMS-1500 02/12 item) | Label | Bounding box (x, y, w, h) | Source layer | Notes |
|---|---|---|---|---|
| Item 1a, insured ID | [REDACTED_ID_01] | 1210, 182, 640, 44 | keyed | pixel-masked, surrogate in label |
| Item 21, ICD indicator | 0 | 1830, 1268, 40, 36 | keyed | ICD-10-CM |
| Item 21A, diagnosis | M54.50 | 132, 1312, 260, 40 | keyed | handwritten |
| Item 24A line 1, dates of service | 03 14 26 / 03 14 26 | 96, 1660, 420, 38 | keyed | |
| Item 24D line 1, CPT/HCPCS + modifier | 97110 GP | 600, 1660, 330, 38 | keyed | shifted half a line |
| Item 24E line 1, diagnosis pointer | AB | 960, 1660, 120, 38 | keyed | |
| Item 24G line 1, units | 2 | 1340, 1660, 90, 38 | adjudicated | keyed value was 4; examiner corrected |
| Item 24J line 1, rendering NPI | [SURROGATE_NPI_07] | 1600, 1660, 300, 38 | keyed | |
| Item 33a, billing NPI | [SURROGATE_NPI_02] | 1180, 2410, 300, 38 | keyed |
A UB-04 record follows the same pattern by form locator: type of bill (FL 4), statement period (FL 6), the revenue-code line grid (FL 42 through FL 47), payer and NPI blocks, principal and other diagnoses (FL 67 series), and attending and operating physician identifiers (FL 76 onward). The revenue-line grid is the hardest region, because line counts, continuation pages and totals interact.
Redacting identifiers in the pixels, not only the metadata
Identifiers on a claim form live in the image, so de-identification must operate on pixels as well as on the keyed record and any file metadata. Names, addresses, member IDs, dates of birth, SSNs on older stock, signatures and account numbers are all printed or handwritten on the page. The DICOM standard makes the same point for imaging: its attribute confidentiality profiles do not guarantee removal of all identifying information and are only one part of a de-identification process [3].
Under HIPAA, health information can be de-identified by Safe Harbor, which removes 18 identifier types, or by Expert Determination [1], with the standard set in 45 CFR 164.514(a)-(b) [2]. A limited data set under 164.514(e) keeps some dates and geography but requires a data use agreement [2]. Dates of service are themselves Safe Harbor identifiers below the year level, which affects how you label Item 24A and FL 6; decide early whether you need Expert Determination to keep real dates.
Pixel redaction has its own failure modes. Black boxes over handwriting can leave ascenders visible, masks drawn from the keyed record miss values written outside their box, and surrogate text rendered in a clean font teaches the model a shortcut. Our guide to redacting PII in scanned documents across pixels, OCR layers and metadata covers these in detail.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Splitting train and evaluation sets for claims intake
Split by submitter and time, not by image, or your evaluation numbers will overstate generalization. Practice-management software prints the same layout, font and alignment offset on every claim from one billing office, so a random image split leaks submitter templates into the test set. Hold out whole billing entities and a later date range.
Score at three levels: field-level exact match after normalization (dates, NPIs, codes), service-line association (did each charge land on the correct line with the correct diagnosis pointer), and claim-level straight-through rate, meaning every required field correct. For outcome-based evaluation across the claim lifecycle, see insurance claims AI evaluation using adjudication outcomes. Selection marks such as Item 1 program type, Item 10 condition flags and Item 27 assignment are a separate error class; checkbox and selection mark data explains how to label them.
Illustrative example: invented to show structure; it does not describe an available dataset.
Request template: scanned claim form dataset
- Forms: CMS-1500 (version per image) and/or UB-04, with continuation pages flagged.
- Strata: dropout original, black-and-white, photocopy, fax; typed, handwritten, mixed; target share for each.
- Ground truth: keyed record per image, source layer per field, QA method and measured keying error rate.
- Geometry: per-field and per-line bounding boxes in a documented coordinate format.
- De-identification: method (Safe Harbor or Expert Determination), pixel redaction approach, surrogate policy, sample check.
- Split metadata: anonymized submitter key and receipt month to support grouped splits.
- Rights: who held the paper and the keyed data, and whether the provider or payer relationship permits licensing.
How these forms differ from neighboring claim documents
CMS-1500 and UB-04 images are the claim as submitted, which separates them from documents produced later in the cycle. Remittance documents are a different task; see paper EOB images paired with posted payments. Dental claims use the ADA form and narratives, covered in dental claims with narratives and attachments. Property and casualty intake uses different standards, covered in ACORD form extraction data.
For general scanned and handwritten forms outside healthcare, the SourceX owner page on licensing scanned forms and handwritten documents and the overview of training data for document understanding apply. Buyers in claims operations can also start from claims administration buyers.
Sourcing CMS-1500 and UB-04 images for claims OCR
SourceX sources operational datasets, including documents and finance workflows, from US companies on request; nothing is held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed, personal details are removed or replaced before delivery with the method recorded and a sample checked, and health records require HIPAA de-identification. Describe the forms, strata and ground truth you need at SourceX for AI data buyers.
Frequently asked questions
Can I synthesize CMS-1500 images instead of licensing real scans?
Synthetic forms help with layout coverage and rare codes, but they rarely reproduce dropout-ink artifacts, fax degradation, handwriting and misregistration as they occur in a real mailroom. Most teams use synthetic data for pretraining and real scans for fine-tuning and evaluation.
Is a public dataset of real claim forms available?
We are not aware of a public, licensed corpus of real CMS-1500 or UB-04 scans with keyed ground truth. Public form benchmarks such as FUNSD cover other domains [4].
Who can license claim form images?
Typically the entity that received and keyed them, such as a payer, TPA or mailroom vendor, subject to its contracts and HIPAA status. Confirm that the holder's agreements permit licensing for model training before negotiating terms.
Sources
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Electronic Code of Federal Regulations, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- NEMA / DICOM Standards Committee, "DICOM PS3.15 Annex E: Attribute Confidentiality Profiles" (2026). https://dicom.nema.org/medical/dicom/current/output/chtml/part15/chapter_E.html
- Jaume, Ekenel, Thiran, "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.