Document AI data
Document Classification Training Data: Taxonomies from Real Business Document Mixes
Quick answer
A useful document classification dataset reproduces your intake stream, not a tidy benchmark. Source labeled pages or files from the same kinds of channels you route (mailroom scans, claims portals, fax inboxes, loan packets), keep the natural class distribution, include an explicit other/unknown class and the long tail, and decide upfront whether labels attach to pages or to logical documents. Public sets such as RVL-CDIP are good for pretraining and smoke tests, but not as a proxy for production accuracy.
By SourceX Editorial · Updated
Why RVL-CDIP is not a production proxy
RVL-CDIP remains the standard document-image classification benchmark, but its age, domain and design limit what it tells you about your own intake. It holds 400,000 scanned document images in 16 classes such as letter, form, invoice and report [1]. It was drawn from the IIT-CDIP collection of tobacco-industry litigation records, so the documents are mostly decades-old typewritten and photocopied corporate paper rather than modern PDFs, portal uploads or phone photos.
Three problems follow for buyers. First, the classes are balanced and closed, so the benchmark never asks a model to reject a document type it has not seen, which is a daily event in a real mailroom. Second, benchmark labels are noisier than they look: Northcutt et al. estimated an average of at least 3.3 percent label errors in the test sets of ten widely used benchmarks [4], so check the label quality of any benchmark, RVL-CDIP included, before relying on leaderboard deltas.
Third, RVL-CDIP is single-page. Van Landeghem et al. argue that page-level benchmarks are outdated for classifying complete documents and note the shortage of public multi-page classification sets [2], and long documents are underrepresented in common classification corpora generally [3]. For a licensing view of public alternatives, see public document AI datasets and commercial use.
Design the taxonomy from the routing decision, not the folder tree
A good classification taxonomy has one class per downstream action: if two document types go to the same extraction model and the same queue, they are probably one class. Start from the routing table in your IDP or BPM system (for example, the document-type field that selects an extraction template) and work backward to classes. Then add three structural classes that benchmarks omit: other/unknown, separator or cover sheet, and unreadable or blank.
Suppliers will arrive with their own labels: ECM folder names, document-type codes from a claims or loan origination system, fax-routing rules, or email subject tags. Those are useful weak labels, but they encode the supplier's workflow, not yours. Agree a written mapping table before delivery, mark many-to-one and one-to-many mappings, and send ambiguous classes to a human review pass. Expect label noise: if curated benchmarks carry measurable test-set errors [4], system-derived labels on business documents are rarely cleaner.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Supplier source label | Source system | Your class | Mapping note | Review? |
|---|---|---|---|---|
CLM_FNOL_FORM | Claims admin system doc-type code | First notice of loss | 1:1 | Sample 2% |
Correspondence | ECM folder | Split: policyholder letter / attorney letter of representation | 1:many, needs relabel | Full review |
MED_BILL, UB04, CMS1500 | Claims attachments | Medical bill | many:1 | Sample 2% |
Misc | ECM folder | Other/unknown (or relabel) | Catch-all, high noise | Full review |
FAX_COVER | Fax server rule | Cover sheet | Page-level only | Sample 5% |
Keep the natural distribution and the long tail
Request the natural class distribution of a defined intake window, plus an oversampled long-tail supplement delivered as a separate, flagged split. Balanced training sets are convenient, but evaluating on them overstates production accuracy because the head classes that dominate real volume are tested no harder than rare ones, and the costly errors in routing usually sit in the tail and in the other class. Keep at least one held-out test split at the true prior so that precision per class and the misroute rate reflect what operators will see.
Ask the supplier for a class-frequency table for the full window before sampling. Note where frequencies are seasonal (open enrollment, quarter-end, catastrophe events for property claims) and capture a window that spans them. If the long tail has fewer than a few dozen examples per class, plan to treat those classes as few-shot or rejection targets rather than training a dedicated head; the golden evaluation dataset guide covers how to freeze such a test set.
Page-level vs document-level labels
Decide the unit of classification before any labeling, because a page-level label and a document-level label answer different routing questions. Mailroom and fax streams usually arrive as page streams in which one TIFF or PDF holds several logical documents, so you need page labels plus document boundaries. Portal uploads usually arrive one file per document, but a single "claim packet" PDF can still contain a form, a bill, photos and a letter.
Specify these rules in the request:
- Unit: page, logical document, or both, with a
doc_idandpage_indexon every page. - Boundaries: first-page flags or split points; see page-stream segmentation data for that task.
- Cover sheets and separators: labeled as their own class at page level, excluded from document-level labels.
- Attachments and exhibits: labeled by their own type with a
parent_doc_id, not by the parent's type. - Multi-type pages: one primary label plus an optional secondary label, never a forced single choice.
For long packets such as mortgage files, the multi-page and long document guide and the mortgage loan file classification page go deeper.
Example intake taxonomies
Real intake taxonomies are narrower than RVL-CDIP and full of near-duplicates that differ only by the action they trigger. Two common shapes:
Healthcare fax inbox. Referrals, orders (DME, imaging, lab), prior authorization requests and responses, records requests, lab and imaging results, discharge summaries, insurance cards, fax cover sheets, and other. Faxes are low-resolution, often skewed and frequently arrive as multi-document streams, so pair this with degraded document image data. These pages carry protected health information, so when the supplier is a HIPAA covered entity or business associate, a licensed set should be de-identified under 45 CFR 164.514 (Safe Harbor or Expert Determination) [5]; detail on health records sits in the industries cluster.
Claims mailroom. First notice of loss, proof of loss, police or incident reports, repair estimates, medical bills, attorney letters of representation, subrogation correspondence, recorded statements transcripts, photos of damage, and other. The hard boundaries are usually letter subtypes, where layout is identical and only content decides the route.
Specify a delivery record that supports audit
Ask for a manifest that ties each label to its source and provenance so you can audit weak labels and reproduce splits. Images alone, or a folder per class, make it impossible to tell which labels came from system metadata and which were human-reviewed.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"doc_id": "d-000412",
"page_index": 3,
"file": "pages/d-000412_p003.png",
"source_format": "TIFF G4 fax, 204x196 dpi",
"channel": "fax_inbox",
"label_page": "lab_result",
"label_doc": "lab_result",
"is_first_page": false,
"label_source": "system_doc_type_mapped",
"source_label_raw": "LAB_RPT",
"human_reviewed": true,
"reviewer_agreement": 1.0,
"split": "test_natural_prior",
"deidentification": "redacted_and_replaced; method logged"
}
Buyer checklist before you sign
Use this list to compare candidate sources on the dimensions that predict production accuracy, not on raw image counts.
- Class-frequency table for the intake window, including other/unknown volume.
- Written mapping from supplier labels to your taxonomy, with ambiguous classes flagged.
- Label provenance per record: system-derived, rule-derived, or human-reviewed, with a reviewed sample and agreement rate.
- Unit of labeling and boundary rules for packets, cover sheets and attachments.
- Capture conditions: native PDF vs scan vs fax vs phone photo, with resolution.
- A held-out split at the natural prior, and no near-duplicate leakage across splits; near-duplicate forms and templates are common in business archives.
- De-identification method and the residual-risk statement for personal data.
- A license that names records, allowed uses (training, evaluation or both), term and delivery.
If you need to label your own archive instead, compare the trade-offs in annotating your own documents vs licensing pre-labeled data. The Document AI data hub lists the related extraction and layout tasks.
How SourceX sources classification data
SourceX sources operational datasets from US companies on request, including documents and the support, finance and legal workflows that produce real intake mixes; datasets are not held in stock, and a request does not guarantee a match. Buyers describe the documents, classes and distribution they need, and SourceX looks for US businesses that hold that data; each release is approved by the supplying company. Every dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phones and account numbers are removed or replaced with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows only after an executed agreement. You can describe a document classification request to SourceX or browse the enterprise document archives overview.
Source a classification set that matches your intake
SourceX sources operational document datasets from US companies on request and manages licensing, with every release approved by the supplying company and delivered under a license that defines records, uses, term and delivery. Describe your taxonomy, units and target distribution, and SourceX assesses data and licensing permissions before anything is agreed. Start a buyer request.
Frequently asked questions
Can I fine-tune on RVL-CDIP and then adapt to my documents?
Yes, as a starting point. Pretraining on RVL-CDIP gives a model generic layout cues, but the closed 16-class taxonomy, single-page format and label noise mean you still need in-domain labeled data and a natural-prior test set to measure real routing accuracy [1][2].
How do VLMs change the data requirement?
Zero-shot vision-language classifiers reduce the training data needed but not the evaluation data. You still need a labeled test set at the real class distribution, with an other class, to measure misroutes and calibrate rejection thresholds; see document extraction evaluation ground truth for the companion extraction test.
Should the other/unknown class be trained or handled by a threshold?
Usually both. Label real out-of-taxonomy documents from the intake stream as other so you can measure rejection, and tune a confidence threshold on the natural-prior split; softmax confidence alone tends to be overconfident on unseen document types, so measure rejection directly.
Sources
- AIKosh (IndiaAI), "RVL-CDIP dataset". https://aikosh.indiaai.gov.in/home/datasets/details/rvl_cdip.html
- arXiv (Van Landeghem et al.), "Beyond Document Page Classification: Design, Datasets, and Challenges" (2023). https://arxiv.org/abs/2308.12896v1
- arXiv, "Long-length Legal Document Classification" (2019). https://arxiv.org/pdf/1912.06905
- arXiv / NeurIPS 2021 (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- eCFR (Office of the Federal Register / HHS), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.