Skip to content

Document AI data

Document Question Answering Data: Questions Grounded in Real Business Documents

Quick answer

A document visual question answering dataset pairs page images with natural-language questions, one or more accepted answer strings and, ideally, the answer's location on the page. Public sets such as DocVQA, InfographicVQA and SlideVQA set the format, and DocVQA made ANLS the standard metric, but they are research benchmarks built on narrow document pools and are now widely trained on. Teams fine-tuning or evaluating document VLMs for commercial use usually need private, licensed page images from real business workflows, with questions written by people who know those documents.

By SourceX Editorial · Updated

What a document VQA record must contain

A usable record holds the page image, the question, every acceptable answer string, an answer type and an evidence location tied to the page's OCR tokens. DocVQA established the pattern of extractive questions over single scanned pages, with 50,000 questions on more than 12,000 images [1]. Commercial sets should go further, because answer boxes and evidence spans are what let you train grounding, audit hallucinations and compute localization metrics.

Store answers as a list, not a single string. "$1,250.00", "1,250.00" and "1250" can all be correct, and ANLS, the metric DocVQA reports, takes the best match across accepted answers [1]. Keep the OCR engine name and version with each page, since a box expressed against Textract word IDs will not line up with Tesseract or Azure Document Intelligence output.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "qa_id": "q-000184",
  "doc_id": "inv-2024-0091",
  "page_index": 0,
  "image": {"file": "inv-2024-0091_p0.png", "dpi": 300, "width": 2550, "height": 3300},
  "question": "What is the balance due after the early-payment discount?",
  "answers": ["$1,212.50", "1,212.50", "1212.50"],
  "answer_type": "arithmetic",
  "answerable": true,
  "evidence": [
    {"bbox": [1840, 2610, 2210, 2660], "ocr_token_ids": [412, 413], "role": "operand"},
    {"bbox": [1840, 2702, 2210, 2750], "ocr_token_ids": [431], "role": "operand"}
  ],
  "reasoning_note": "Total 1,250.00 minus 3% discount listed in terms block",
  "ocr": {"engine": "example-ocr", "version": "x.y"},
  "annotator": {"role": "AP specialist", "guideline": "dqa-guide-v3", "pass": "double"},
  "split": "eval-private",
  "redaction": {"method": "replace-synthetic", "fields": ["payee_name", "account_no"]}
}

Coordinates in pixel space against a declared resolution avoid the silent errors that normalized 0-1000 boxes cause when images are later resized. If the answer is computed rather than read, as above, record every operand box rather than inventing a box for a number that never appears on the page.

Question taxonomy and balance

A balanced set mixes direct lookups, arithmetic, yes/no, comparison, multi-hop and unanswerable questions in proportions you set before annotation starts. DocVQA found that models trail humans (94.36%) most on questions that depend on document structure such as tables, forms and layout [1].

Left alone, annotators drift toward the easiest question type: copy a visible key-value pair. That inflates scores and teaches little. A quota per type, enforced at batch acceptance, is the simplest control.

Illustrative example: invented to show structure; it does not describe an available dataset.

Question typeExample on a purchase orderSuggested share (training)Evidence required
Extractive lookup"What is the PO number?"30-40%One box
Table cell lookup"What unit price is listed for line 3?"15-20%Cell box plus row header
Arithmetic"How many units across all lines?"10-15%All operand boxes
Yes/no or verification"Is the PO signed?"5-10%Signature region or absence note
Multi-hop / cross-region"Which vendor ships the item with the latest date?"10-15%Boxes for each hop
Multi-page"Does the total on page 3 match the summary on page 1?"5-10%Page index plus boxes
Unanswerable"What is the buyer's tax ID?" (not present)5-10%Explicit null answer

Unanswerable questions are the cheapest guard against confident hallucination, and most public single-page sets underrepresent them. Write them against the same page so the model cannot learn that an unfamiliar layout implies "no answer."

Answer localization and evidence

Answer boxes turn a QA set into grounding supervision, so specify their granularity, coordinate system and relation to OCR tokens in the statement of work. Token-level evidence (OCR word IDs) is more reusable than free-drawn rectangles because it survives re-rendering and lets you compute span overlap. For answers that require reasoning, ask for operand boxes plus a short reasoning note, not a single box around the result.

Multi-page evidence needs a page index on every box. MP-DocVQA was built because earlier DocVQA benchmarks were single-page and long inputs strain transformer models [3], and DUDE extends evaluation to multi-page documents with varied answer types [4]. If your target is long contracts or loan packets, see multi-page and long document data for packet-level specifications.

ANLS and what else to measure

ANLS scores each prediction as one minus the normalized edit distance to the closest accepted answer, sets scores below 0.5 to zero and averages over questions; DocVQA adopted it as its headline metric [1]. It tolerates OCR noise such as a dropped comma, which is why DocVQA-family leaderboards report it. It also means a wrong but similar number, such as 1,215.50 against 1,212.50, still earns partial credit.

For business documents, report ANLS alongside exact match on normalized values (amounts, dates, IDs) so numeric errors are not hidden. Where answers are lists, score each item separately and penalize missing or extra entries. Add evidence IoU or token-overlap against the gold boxes, and abstention precision on unanswerable questions; together these catch models that are right for the wrong reason.

Why public DocVQA-style sets fall short for commercial work

Public document VQA sets are narrow in document mix, are released under research-oriented terms you must check one by one, and are likely present in many pretraining corpora. DocVQA pages come from the UCSF Industry Documents Library [1], and DuReader-vis draws on images from Baidu search [5]. None of those pools look like a mid-market company's invoices, claims files or engineering change orders.

Contamination matters most for evaluation. If a benchmark's images and questions have circulated since 2020, high scores may reflect memorization rather than reading. A private, held-out set built on documents that have never been public is the cleaner test; the document extraction evaluation guide covers how to keep such a set sealed. For "DocVQA commercial use" questions, review each dataset's terms on the distribution portal before shipping a model trained on it; our public document datasets and commercial use page compares the common ones.

"InfographicVQA alternative" searches usually come from teams whose real documents are charts, dashboards and slide decks. SlideVQA's 52K slide images and 14.5K multi-hop and numerical questions show the format [2], but internal decks and reports carry the domain vocabulary your users will ask about.

Who writes the questions, and under which guideline

Question realism depends on who writes them, so record annotator role, guideline version and review pass on every record. Domain staff such as AP clerks, claims adjusters or paralegals ask the questions their colleagues actually ask; generalist annotators tend to paraphrase visible labels. Crowd-written questions can also cluster in the most prominent regions of a page, which you can check by plotting answer-box centroids.

Specify a written guideline before collection: allowed question types, banned patterns (questions answerable without the image, questions that quote the answer's label verbatim), answer normalization rules and how to handle illegible text. Double annotation on a sample, with disagreements adjudicated, gives you an agreement figure to compare against model scores. LLM-generated questions can pad training splits, but keep evaluation questions human-written and see synthetic vs real documents for where generated data breaks.

Sourcing specification and acceptance checklist

A sourcing brief for document VQA data should fix the document mix, question quotas, evidence format, split policy, redaction method and license scope before any supplier starts work. The checklist below doubles as acceptance criteria for each delivered batch.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Document mix: named document types (invoices, POs, bank statements, lab reports, change orders), share of scans vs born-digital PDFs, and capture conditions; see degraded document images.
  • Page format: PNG or TIFF at a stated DPI, plus the source PDF where licensed; OCR JSON per page with engine and version.
  • Question quotas: per-type shares as in the table above, with at least some unanswerable and multi-page items.
  • Answers: list of accepted strings, normalized value field, answer type, answerable flag.
  • Evidence: pixel-space boxes linked to OCR token IDs; page index for multi-page items.
  • Provenance: annotator role, guideline version, review pass, date.
  • Splits: document-level splits so no page appears in both train and eval; a sealed evaluation split.
  • Privacy: how names, account numbers and addresses were removed or replaced, and whether answers that were personal data were dropped or substituted consistently.
  • Rights: a license that covers the page images and the derived question-answer pairs and boxes together, for the uses you intend.
  • Acceptance: sample re-annotation agreement, ANLS of a reference model on a held-out slice, and box-to-token alignment rate.

Redaction interacts with QA in a specific way. If a payee name is replaced on the image but not in the answer list, the record is broken; replacements must be applied consistently to pixels, OCR text and answers. Health documents add HIPAA de-identification requirements under 45 CFR 164.514 [6].

How SourceX fits a document VQA project

SourceX sources operational datasets from US companies on request, including documents and finance and legal workflow records, and manages the licensing and ongoing purchase process. Nothing is held in stock, and a request does not guarantee a match. You describe the documents you need; SourceX looks for US businesses that hold that data, and each release is approved by the supplying company.

Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery, which is where coverage of derived QA pairs belongs. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows after an executed agreement. SourceX does not train models. You can start a document VQA data request with your question taxonomy and evidence format, and read more on training data for multimodal document models.

If your use case is passage retrieval over corpora rather than page-image QA, the RAG evaluation datasets from company documents page and retrieval-augmented fine-tuning data are the better fit. For the full map of document tasks and annotation layers, start at the document AI datasets hub.

Source licensed document VQA data for your model

SourceX sources operational documents from US companies on request and manages licensing, with personal details removed or replaced before delivery and every release approved by the supplier. Describe your document types, question taxonomy and evidence format to begin. Submit a buyer request.

Frequently asked questions

Can I use DocVQA to train a commercial model?

Check the terms on the dataset's distribution portal and the rights in the underlying images before relying on it. The DocVQA images come from an archive of industry documents [1], so both the annotation terms and the source collection matter. Counsel should review the answer for your specific use.

Do I need answer bounding boxes if my model only outputs text?

Boxes are still worth collecting. They let you check whether a correct answer came from the right region, train grounding later without re-annotating and debug failures by comparing predicted and gold evidence.

Is ANLS enough to compare document VQA models?

No. ANLS gives partial credit to near-miss strings, including wrong numbers [1], so pair it with exact match on normalized values, evidence overlap and abstention metrics on unanswerable questions.

How many questions per page should a training set have?

There is no fixed number. Several varied questions per page, across types, usually teach more than one question each on many near-identical pages, but deduplicate templates and split by document so evaluation stays clean.

Sources

  1. arXiv (Mathew, Karatzas, Jawahar), "DocVQA: A Dataset for VQA on Document Images" (2020). https://arxiv.org/pdf/2007.00398
  2. arXiv (Tanaka et al.; AAAI 2023), "SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images" (2023). https://arxiv.org/pdf/2301.04883
  3. alphaXiv overview of arXiv:2212.05935, "Hierarchical multimodal transformers for multi-page DocVQA (MP-DocVQA)" (2022). https://alphaxiv.org/overview/2212.05935v2
  4. ICCV 2023 (CVF Open Access), "Document Understanding Dataset and Evaluation (DUDE)" (2023). https://openaccess.thecvf.com/content/ICCV2023/html/Van_Landeghem_Document_Understanding_Dataset_and_Evaluation_DUDE_ICCV_2023_paper.html
  5. ACL Anthology (Findings of ACL 2022), "DuReader-vis: A Chinese Dataset for Open-domain Document Visual Question Answering" (2022). https://preview.aclanthology.org/setup/2022.findings-acl.105
  6. eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data