Document AI data
Real-World Document Capture Conditions: Degraded Scans, Faxes, Photocopies and Phone Photos
Quick answer
A useful degraded document images dataset holds real business pages captured the way production traffic arrives: low-DPI scans, Group 3 faxes, multi-generation photocopies and phone photos, each labeled per page with its capture conditions. Public sets are mostly historical, handwritten or synthetic, so they rarely match modern intake streams. Buyers should specify a condition taxonomy, demand per-page condition labels and capture metadata, and keep a real held-out set to test whether synthetic augmentation actually closes the gap.
By SourceX Editorial · Updated
Why models trained on clean scans fail in production
Models fail on production documents because their training distribution was clean, born-digital or synthetically degraded, while intake queues are dominated by artifacts the model never saw. The well-known public reference for noisy forms, FUNSD, contains only 199 noisy, low-resolution scanned forms [1], which is enough to benchmark but not enough to represent a claims, lending or freight intake stream. Research on document image classifiers also shows that scores on an in-distribution test split say little about performance on out-of-distribution documents [2].
The typical failure modes are specific. OCR can drop characters on normal-mode fax text, layout models can merge columns on skewed pages, key-value extractors may read a highlighter stroke as a field boundary, and classifiers confuse a faxed cover sheet with the first page of the packet. Each of these is a capture problem, not a document-type problem, which is why capture conditions deserve their own labels across every type covered in the Document AI datasets by task hub.
A capture-condition taxonomy that buyers can label against
Start with a fixed taxonomy so suppliers, annotators and evaluators describe degradation the same way. Conditions cluster into five families, and a single page often carries several at once.
- Resolution and sampling: effective DPI, bitonal vs grayscale vs color, JPEG or JBIG2 compression artifacts, downsampling from upstream systems.
- Geometry: skew, rotation, perspective distortion, page curl, folds, cropped margins, partial pages.
- Photometric: blur (motion and defocus), glare, shadows, uneven illumination, low contrast, toner fade.
- Physical media: bleed-through, show-through, stains, punch holes, staples, multi-generation photocopy noise and edge darkening.
- Overlays: stamps, signatures, handwritten annotations, highlighter, sticky notes and redaction boxes over printed text.
Fax and phone capture deserve their own top-level flags because they combine several families at once. Keep the taxonomy versioned; when it changes, relabel the evaluation set before comparing models.
Fax pages: low-resolution bitonal images with structural noise
Fax images are a distinct capture condition because the transmission standard fixes their resolution and color depth. Group 3 fax pages in normal mode carry roughly half the vertical resolution of fine mode, and they are bitonal, so anti-aliasing and gray levels are gone before OCR starts. Normal-mode pages have anisotropic pixels, which distorts character aspect ratios unless the pipeline resamples correctly.
Faxes also add structure that models must learn to ignore or use: transmission header lines with sender numbers and timestamps, cover sheets, page-count footers and repeated noise from line errors. Healthcare and legal intake streams have historically been fax-heavy, so faxed records often contain protected health information. Health records require HIPAA de-identification under Safe Harbor or Expert Determination before release [6], and header lines with fax numbers are an easy-to-miss identifier.
Phone photos: perspective, curl, glare and backgrounds
Phone-captured documents fail models through geometry and lighting rather than resolution. A typical capture shows perspective distortion, page curl, glare on glossy stock, shadows from the photographer's hand, and a cluttered background such as a desk, steering wheel or car seat. Document boundary detection, perspective correction and dewarping all sit before OCR in a mobile pipeline, and each needs its own evaluation.
Dewarping training data is hard to source from operational archives, because supervised dewarping needs a flat reference image or a deformation field for every capture; research sets commonly fill the gap by synthetically warping flat pages. Business archives hold the captured photo but almost never the flat reference, so licensed phone photos mainly serve end-to-end evaluation of capture plus extraction rather than supervised dewarping. Public camera-capture sets are often narrow; one IEEE DataPort set of calligraphy photos has a 100-image test split under uniform and random lighting [4]. Verify the license terms of any public dewarping set before commercial use, and see which public document datasets allow commercial use.
Phone photos also capture incidental personal data: faces, hands, other documents on the desk, screens and license plates. Screen backgrounds as well as the document itself before any release.
Real vs augmented degradation: where synthetic artifacts break
Synthetic augmentation reproduces geometric and blur artifacts reasonably well and reproduces fax dithering, real bleed-through and multi-generation photocopy noise poorly. Augmentation libraries apply noise to a clean image independently of content, while real bleed-through mirrors the text on the reverse side and real photocopy noise compounds across generations with toner and platen effects.
The practical test is simple: train with augmentation, then evaluate on a real held-out set stratified by condition. If error on real fax or photocopy strata stays well above error on augmented equivalents, the augmentation is not covering that condition and real pages are needed. The broader trade-offs are covered in where synthetic documents break.
Per-page condition labels and image quality scores
Label conditions per page, not per document, so evaluation can be stratified and quality gates can be set at the page level. A multi-page packet can mix a clean typed page, a faxed lab result and a phone photo of a signature page, and a document-level label hides which page caused the failure. Pair the labels with transcriptions from OCR ground truth aligned to real scans so character error rate can be broken out by condition.
For document image quality assessment, add a graded quality score. MHDID provides 335 historical document images in four distortion classes with human opinion scores [3], a useful model for structure even though the content is historical rather than business. Record labels in machine-readable metadata; Croissant-RAI is one vocabulary for documentation such as provenance and labeling [5].
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"page_id": "pkt-000412-p03",
"packet_id": "pkt-000412",
"document_type": "lab_result",
"capture": {
"channel": "fax",
"fax_mode": "normal",
"dpi_x": 204,
"dpi_y": 98,
"color_depth": "bitonal",
"file_format": "TIFF",
"compression": "CCITT_G3",
"device_metadata_available": false
},
"conditions": ["fax_header_line", "skew", "toner_fade", "stamp_over_text"],
"skew_degrees": 2.4,
"quality_score_1_to_5": 2,
"quality_rater_count": 3,
"has_ocr_ground_truth": true,
"deidentification": {"method": "replaced", "fields": ["patient_name", "fax_number"]}
}
Capture metadata and DPI requirements to put in the request
Ask suppliers for scanner and capture-setting metadata where it exists, because it lets you explain failures instead of guessing. Useful fields include EXIF data from phone captures (device model, focal length, exposure), TIFF tags from scanners (XResolution, YResolution, Compression, PhotometricInterpretation), the capture application, and whether the image was re-encoded by a document management system. Many archives strip this on ingest, so record its absence explicitly rather than inferring DPI from pixel dimensions.
DPI requirements depend on the target, not on a universal number. If production includes normal-mode fax and downsampled email attachments, the training and evaluation mix must include them at native resolution; upsampling everything to 300 dpi in preprocessing hides the condition the model must handle. Put the expected mix of conditions in the request, as described in the document dataset requirements spec.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Condition stratum | Target share of eval set | Must include | Acceptance check |
|---|---|---|---|
| Flatbed or ADF scan, 200-300 dpi | 30% | Skew, punch holes, stamps | Native TIFF or PDF, TIFF tags retained |
| Group 3 fax, normal and fine | 20% | Header lines, cover sheets | Bitonal, unresampled |
| Multi-generation photocopy | 15% | Edge darkening, toner fade | Generation count noted where known |
| Phone photo | 25% | Perspective, curl, glare, background | EXIF retained or absence noted; backgrounds screened |
| Mixed overlays | 10% | Highlighter, handwriting over print | Overlay type labeled per page |
How SourceX approaches capture-condition data requests
SourceX sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases; documents are among the kinds of data it sources. Data is sourced on request rather than held in stock, so a request for faxed intake pages or phone-captured forms is a description of what you need, not a catalog pick, and it does not guarantee a match. Buyers describe the data, not the businesses; SourceX looks for US businesses that hold it, and every release is approved by the supplying company.
Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe the capture conditions you need, and related scanned material is covered on the scanned forms and handwritten documents page and in the OCR glossary entry. For non-document photos, see image datasets for computer vision.
Request degraded document images for your capture pipeline
SourceX sources document data from US companies on request, rights-reviewed and delivered under a license that defines records, uses, term and delivery. Nothing is contracted until a supplier agrees, and delivery runs through private, access-controlled workflows after an executed agreement. Describe the capture conditions, document types and labels you need at SourceX for buyers.
Frequently asked questions
Can public datasets cover degraded business documents?
Partly. FUNSD offers 199 noisy scanned forms [1] and MHDID offers quality-scored historical pages [3], but neither represents a modern mix of faxes, photocopies and phone photos at scale, and license terms vary by set.
Should phone-captured documents be used for dewarping training?
Usually only for evaluation. Supervised dewarping needs a flat reference or deformation ground truth for each capture, and operational archives rarely hold it.
How many pages per condition are needed?
There is no fixed number. Size each stratum so the confidence interval on your key metric, such as field-level F1 or character error rate, is narrow enough to compare two pipelines, and add pages to strata where results swing between runs.
Sources
- arXiv (Jaume, Ekenel, Thiran), "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
- NeurIPS 2022 Datasets and Benchmarks (Larson et al.), "Evaluating Out-of-Distribution Performance on Document Image Classifiers" (2022). https://proceedings.neurips.cc/paper_files/paper/2022/hash/4c0986bd04d747745beba3752bdf4d9d-Abstract.html
- Qatar University QSpace, "MHDID: A Multi-distortion Historical Document Image Database". https://qspace.qu.edu.qa/handle/10576/13069
- IEEE DataPort, "Degraded document images". https://ieee-dataport.org/documents/degraded-document-images
- arXiv (Jain et al., MLCommons Croissant RAI task force), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
- eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.