Skip to content

Document AI data

Multi-Page and Long Document Data for Document Understanding Models

Quick answer

A multi-page document understanding dataset pairs whole documents, typically tens to hundreds of pages, with labels and questions whose answers depend on more than one page: a table that continues across a page break, a defined term used 40 pages after its definition, an exhibit referenced from the body, or a total that only appears on the last page. Useful sets report page and token length distributions, cite evidence pages and regions for every answer, and ship native PDFs alongside page images and per-page OCR.

By SourceX Editorial · Updated

Most public sets were built for research and skew short or academic, so teams evaluating long-context vision language models (VLMs) usually need to check length, domain and license before assuming coverage. This guide covers what to specify, how to label for grounding, and how to audit a long-document set. For the broader map of document tasks, start at the Document AI data hub.

Why single-page document sets miss long-document failures

Single-page benchmarks cannot test whether a model finds, connects and reconciles evidence spread across a document. MP-DocVQA was created precisely because DocVQA questions sit on one page, and transformer models struggle with long inputs because attention cost grows quadratically with sequence length [1]. Classic form sets such as FUNSD are one-page scans, so strong scores there say little about a 90-page credit agreement.

The failure modes that matter in production are document-level:

  • Continued tables. A schedule of fees or a loan amortization table spans pages 12-15; the header row appears only on page 12, and models misassign columns on later pages.
  • Distant definitions. "Material Adverse Effect" is defined in section 1.1 and applied in section 8.4; the answer requires both.
  • Exhibit and footnote references. "See Exhibit C" or "Note 14" sends the reader to an appendix that may be a scanned attachment with a different layout.
  • Rollups. Subtotals on interior pages must reconcile with a grand total on the final page.
  • Superseded content. Amendments, restated sections or later-dated riders override earlier text in the same file.

These phenomena also separate this task from page-stream segmentation, which splits a batch of concatenated, separate documents. Here the unit is one long, coherent document.

What public multi-page benchmarks cover, as of October 2026

Public long-document sets are useful baselines, but their length, domain and question mix rarely match an enterprise workload. As of October 2026, the main reference points are:

SetWhat it offersWatch for
MP-DocVQA [1]Multi-page extension of DocVQA questionsAnswers often still sit on a single page; documents are short relative to contracts or filings
DUDE [2]Multi-industry, multi-domain, multi-page visually rich documents with multi-task evaluationCheck the page-count distribution and license before commercial use
SlideVQA [3]Questions over multiple slide imagesSlide decks, not dense prose or tables
Doc-750K [4]Roughly 758K questions over 3.1M images with cross-page dependenciesAcademic-domain sources; little business-document coverage

Two practical conclusions follow. First, research multi-page QA sets tend to be short or dominated by single-page evidence, so compute your own histogram of pages per document and evidence pages per question before treating a set as long-context coverage. Second, OCR quality and image resolution interact with context length, so buyers should keep both a text path (OCR plus a language model) and an image path in any evaluation to see which one actually fails. The public document datasets for commercial use page covers license checks for these sets.

How to label long documents so graders can verify grounding

Every label in a long-document set should cite the page numbers and regions that justify it, because otherwise a correct answer and a lucky guess look identical. Evidence-page supervision is also becoming a training signal in its own right; DocR1 uses evidence pages to guide reinforcement learning for multi-page understanding [5].

A workable label record has four layers:

  1. Document-level fields: values that summarize the whole file (governing law, effective date, total facility amount, fiscal year end), each with a list of evidence spans.
  2. Cross-page structures: table objects with a table_id that persists across pages, plus continues_from and continues_to pointers.
  3. Questions: typed as single-page, cross-page, aggregation, or unanswerable, with evidence spans and answer format.
  4. Page-level layers: layout boxes, reading order and OCR tokens per page, which make the higher layers checkable. See reading order data for multi-column pages.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "doc_id": "doc_000412",
  "source_format": "pdf_native",
  "page_count": 118,
  "text_tokens_cl100k": 64210,
  "pages_scanned": [97, 98, 99, 100, 101, 102],
  "tables": [
    {"table_id": "t_07", "title": "Schedule of Fees",
     "pages": [12, 13, 14, 15], "header_page": 12,
     "regions": [{"page": 12, "bbox": [72, 188, 540, 742]},
                 {"page": 15, "bbox": [72, 60, 540, 410]}]}
  ],
  "questions": [
    {"q_id": "q_031", "type": "cross_page",
     "question": "Does the fee in Schedule row 18 change after the first amendment?",
     "answer": "Yes, from 0.25% to 0.30%",
     "evidence": [{"page": 14, "bbox": [72, 512, 540, 530]},
                  {"page": 99, "bbox": [90, 220, 520, 300]}]},
    {"q_id": "q_032", "type": "unanswerable",
     "question": "What is the prepayment penalty for year 6?",
     "answer": null, "evidence": []}
  ]
}

Coordinates should state their unit and origin (PDF points from bottom-left, or pixels from top-left at a stated DPI), because mixed conventions silently break region grounding when pages are re-rendered.

Length distributions drive context-window planning

Report length as a distribution in both pages and tokens, because the long tail, not the mean, decides whether a model needs chunking, retrieval or a larger context window. A set with a mean of 40 pages and a 95th percentile of 300 pages behaves very differently from one capped at 60.

Ask for, or compute, these statistics per split:

  • Pages per document: min, median, p90, p99, max.
  • Text tokens per document with a named tokenizer, plus image tokens per page at the resolution you plan to use.
  • Evidence distance: page gap between the earliest and latest evidence for each cross-page question.
  • Share of scanned versus born-digital pages, and pages with tables, charts or handwriting.
  • Questions per document, so a few very long files do not dominate the score.

The long-context evaluation on real documents page covers how to turn these distributions into evaluation buckets beyond needle-in-a-haystack tests.

Delivery formats both text and vision pipelines can use

Ask for native PDFs plus rendered page images and per-page OCR so text-only, vision-only and hybrid pipelines can run on the same documents. Native PDFs preserve the text layer, fonts and object structure; page images (PNG or TIFF at a stated DPI) feed VLMs directly; per-page OCR output (for example hOCR, ALTO XML or JSON with word boxes) lets you compare OCR-plus-LLM baselines against end-to-end models.

Keep identifiers stable across all three: doc_id, 1-based page_index, and the PDF page label where it differs ("iv", "A-3"). Mismatched page numbering between the PDF viewer and the label file is one of the most common causes of grounding errors. Teams training parsers should also see PDF parsing training data for page-to-Markdown targets.

Real versus synthetic long documents

Synthetic generation can scale question counts on long documents, but real business files supply the structural irregularity that models fail on. NVIDIA's developer note on training a VLM for long documents describes an iterative synthetic data generation approach [6], which is a reasonable way to bootstrap training questions over documents you already hold.

What synthesis rarely reproduces is the mess of real files: amendments stapled after signature pages, scanned exhibits inside born-digital PDFs, inconsistent table headers, and redaction boxes. Use synthetic questions for training volume and keep a held-out evaluation set built on real documents with human-verified evidence. Real long documents of interest include credit agreements, annual reports and financial statements, engineering specifications, policy manuals and contract sets such as those described on the contract redline datasets page.

Quality checks before you train or report scores

Audit long-document labels on a sample before trusting any score, because errors compound when one answer depends on several pages. Test-set label errors averaging at least 3.3% have been shown to destabilize benchmark rankings on widely used datasets [7], and long-document questions have more places to go wrong.

Illustrative example: invented to show structure; it does not describe an available dataset.

Long-document acceptance checklist

  • Page-count histogram matches the requested distribution, including the p90 tail.
  • At least the agreed share of questions has evidence on two or more pages.
  • Every evidence region renders on the cited page at the delivered DPI.
  • Continued tables share one table_id; header propagation is checked on a sample.
  • Unanswerable questions are verified absent from the whole document, not just the cited pages.
  • No document appears in both train and test, including near-duplicate versions or amendments.
  • Personal details are removed or replaced, and the method is recorded.
  • A dataset card records license, sources, size and known gaps [9].

ISO/IEC 5259-2 gives a vocabulary of measurable data quality characteristics, such as completeness and accuracy, for reporting these checks consistently [8]. The document dataset requirements spec turns this checklist into a request template.

Sourcing long business documents through SourceX

SourceX sources operational datasets, including documents and finance and legal workflows, from US companies on request; it does not hold stock, and a request does not guarantee a match. Buyers describe the documents they need, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery.

Before delivery, personal details such as names, emails, phone numbers and account numbers are removed or replaced, the method is recorded and a sample is checked, though no method is perfect. That matters for long documents, where identifiers recur across signature blocks, schedules and exhibits. You can describe page ranges, document types and label needs on the SourceX buyer request page, or browse related enterprise document archive datasets and document understanding training data.

Request multi-page document data for long-context models

SourceX looks for US companies that hold the long documents you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered through private, access-controlled workflows. Nothing is contracted until a supplier agrees. Describe your page ranges, document types and label layers at https://sourcex.si/buyers.

Frequently asked questions

How many pages should a long-document evaluation set include?

Match your production distribution rather than a fixed number. If production files reach 200 pages, include a tail bucket at that length, and keep enough questions per length bucket to compare scores across buckets.

Should unanswerable questions be part of a multi-page set?

Yes. Unanswerable questions measure hallucination over long inputs, where a model is tempted to answer from a nearby but wrong section. Verify absence against the full document, not only the pages a reviewer expected.

Can document-level labels come from systems of record?

Often. Totals, dates and counterparties stored in an ERP or contract system can serve as weak labels for document-level fields, but they need evidence regions added by annotators to support grounding checks.

Sources

  1. alphaXiv overview, "Hierarchical multimodal transformers for Multi-Page DocVQA (MP-DocVQA)" (2022). https://alphaxiv.org/overview/2212.05935v2
  2. ICCV 2023 (CVF Open Access), "Document Understanding Dataset and Evaluation (DUDE)" (2023). https://openaccess.thecvf.com/content/ICCV2023/html/Van_Landeghem_Document_Understanding_Dataset_and_Evaluation_DUDE_ICCV_2023_paper.html
  3. arXiv (AAAI 2023), "SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images" (2023). https://arxiv.org/pdf/2301.04883
  4. Emergent Mind, "Doc-750K". https://www.emergentmind.com/topics/doc-750k
  5. arXiv, "DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding" (2025). https://arxiv.org/pdf/2508.07313
  6. NVIDIA NeMo Data Designer documentation, "Training a VLM to Understand Long Documents: An Iterative SDG Story". https://docs.nvidia.com/nemo/datadesigner/latest/dev-notes/vlm-long-document-understanding
  7. arXiv (NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  8. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Data quality for analytics and ML, Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
  9. Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data