Multimodal and embodied data
Evaluating Multimodal RAG on Real Enterprise Content: Slides, Diagrams and Screenshots
Quick answer
A multimodal RAG evaluation dataset pairs a frozen corpus of real files (decks, PDFs with charts, architecture diagrams, UI screenshots) with questions, reference answers and page-level evidence citations, and tags each question by where its evidence lives. Public benchmarks cover document VQA and text-heavy enterprise RAG, but mostly on fictional or public content. To know whether your assistant reads your company's slides and screenshots, you need a held-out set built from content that looks like yours, scored on both citation and answer, per modality.
By SourceX Editorial · Updated
Why text-centric RAG benchmarks miss visual evidence
Most enterprise RAG benchmarks test whether a system finds and quotes text, not whether it reads a bar chart or a swimlane diagram. RAG-Multi-Corpus, for example, mixes PDF, Markdown, HTML, DOCX and PPTX files and provides 786 query-answer pairs with ground-truth citations, but its 236 documents describe five fictional organizations [1]. WixQA grounds questions in one company's support knowledge base, which is mostly article text [3]. These are useful for chunking and retrieval regressions; they say little about a revenue figure that exists only as a chart label in slide 14.
Document VQA research shows why the gap matters. DocVQA's 50,000 questions on 12,000+ document images found models furthest behind humans (94.36% accuracy) on questions that depend on layout and structure [4]. SlideVQA requires reasoning across several slide images in one deck [5], and MMLongBench-Doc places evidence in charts, tables and images inside long PDFs, with 1,082 expert-annotated questions over 135 documents in its NeurIPS 2024 version [6]. None of these is a RAG benchmark over a private enterprise corpus, which is the setting your system will actually run in.
For text-only failure modes such as stale or contradictory versions, see our guide to evaluating RAG on outdated and conflicting documents; this page focuses on evidence that is not text.
What a usable multimodal RAG eval record contains
A usable record ties each question to an exact file, page or frame, and region, so the grader can check retrieval and answer independently. EnterpriseRAG-Bench frames enterprise evaluation around company internal knowledge and cited documents, which lets retrieval failure be separated from generation failure [2]. For visual evidence, a document ID alone is too coarse: a 60-slide deck may contain three charts with similar titles.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"qid": "q-0412",
"question": "Which region missed its Q3 bookings target by the largest margin?",
"reference_answer": "EMEA, about 18% below target",
"answer_type": "extractive_numeric",
"evidence": [
{
"doc_id": "deck-ops-review-2025q3.pptx",
"page_or_slide": 14,
"bbox": [112, 220, 860, 610],
"modality": "chart",
"render": "slide-14.png@150dpi"
}
],
"text_layer_contains_answer": false,
"requires_multi_hop": false,
"acceptable_variants": ["EMEA", "Europe, Middle East and Africa"],
"unanswerable": false,
"corpus_snapshot": "2026-09-30",
"rights_status": "approved_for_internal_eval",
"redaction_applied": ["names", "customer_logos"]
}
Two fields do most of the diagnostic work. text_layer_contains_answer tells you whether a text-only pipeline (OCR plus chunking) could have answered at all, so you can split "visual reasoning" failures from "extraction" failures. render pins the exact rasterization, because a chart re-rendered at a different DPI or from a different PowerPoint version can change what OCR and vision encoders see.
Tagging evidence by modality and scoring per tag
Tag every question with the modality its evidence sits in and report metrics per tag; an aggregate score hides the slice your users complain about. A system that scores well overall can still fail every question whose answer is an arrow in an architecture diagram. Keep the taxonomy short and mutually exclusive so annotators agree.
| Modality tag | Typical source | Common failure mode | Suggested answer metric |
|---|---|---|---|
chart | Bar, line and pie charts in decks and BI exports | Misread axis scale, legend color swapped, value interpolated | Numeric tolerance (for example plus or minus 2%) plus exact entity |
table_image | Tables pasted as images, scanned statements | Row/column misalignment, merged cells dropped | Exact match on cell value; ANLS for strings |
diagram | Architecture, flow, org and swimlane diagrams | Edge direction lost, nodes read as a list | Graded rubric on relation (A calls B, not B calls A) |
screenshot | UI screenshots in tickets, runbooks, wikis | Wrong field read, state (disabled, error) ignored | Exact match on field or status text |
slide_layout | Slide where meaning depends on position or grouping | Callout attached to wrong item | Rubric with explicit position cue |
text | Native text layer (control group) | Chunk boundary, distractor passage | Exact match or LLM-judge with reference |
For short string answers read from images, ANLS (Average Normalized Levenshtein Similarity) is the standard soft metric, introduced because hard accuracy unfairly penalizes small OCR-induced mismatches [7]. Use it for screenshot and table-image strings, numeric tolerance for chart values, and a written rubric for diagram relations. Keep a text control slice so you can tell whether a model upgrade improved visual reading or just general answering.
Scoring retrieval and answers together with cited evidence
Require the system to return cited evidence IDs and grade the citation before the answer; a correct answer with the wrong citation is a grounding failure. This is the same principle as requiring cited document IDs in enterprise RAG benchmarks [2], applied at page and region granularity.
A practical scoring stack has four numbers per question:
- Evidence recall@k: did any of the top-k retrieved units (page images, slide renders, crops) contain the gold evidence page or slide?
- Citation precision: of the citations the answer emits, how many point to gold evidence? Penalize citing a whole 60-slide deck.
- Answer correctness: per-modality metric from the table above.
- Grounded correctness: answer correct and at least one citation correct. This is the headline number to report.
Also include unanswerable questions, where the corpus genuinely lacks the fact, and score abstention. Visual RAG systems can hallucinate plausible chart values when the right slide is not retrieved, and unanswerable items expose that.
Building the corpus from real enterprise files
The corpus should be a frozen snapshot of real file types in their native formats, rendered once, with rights and redaction recorded per file. Synthetic decks generated from templates under-represent the mess that breaks production systems: slides with charts embedded as images rather than native chart objects, Visio exports flattened to PNG, screenshots with cropped toolbars, and scanned PDFs without a text layer.
A workable build sequence:
- Collect native files (PPTX, PDF, DOCX, XLSX exports, PNG/JPEG screenshots, SVG or Visio-derived diagrams) plus their metadata: owner system, created and modified dates, and version.
- Render each page or slide to a fixed image (record DPI and renderer) and keep the native text layer separately so you can compute
text_layer_contains_answer. - Author questions from people who know the domain, written against a specific region, with a second annotator confirming the answer and evidence. Target a balanced count per modality tag rather than whatever the corpus happens to contain.
- Add distractors: near-duplicate decks from adjacent quarters and similar dashboards, so retrieval has to discriminate.
- Freeze and version the snapshot; date every question.
Screenshot-heavy sources such as support tickets are a good fit for the screenshot slice; see support tickets with screenshots and attachments. For deck-specific sourcing, the presentation deck datasets page covers what slide corpora contain, and RAG evaluation datasets from real company documents covers the text-centric case.
Rights, redaction and confidentiality in visual evidence
Real decks and diagrams carry confidential business data and third-party content, so they need rights review and redaction before they enter any eval set. A single slide can contain a customer logo, a partner's pricing, a stock photo under a separate license, and an employee headshot. Screenshots routinely show names, email addresses and account numbers in UI fields that no text-layer scrubber will catch.
Practical checks for buyers:
- Ownership per element: is embedded imagery, a third-party chart or a pasted analyst figure licensable by the company supplying the deck? See licensing multimodal records with several rightsholders.
- Pixel-level de-identification: redaction must happen on the rendered image, not only the text layer, and must not destroy the evidence region. See de-identifying multimodal records.
- Permitted use: an evaluation-only grant, a grounding license and a training license are different permissions; grounding vs training licenses explains the split.
- Redaction-evidence conflict: if redaction blacks out the answer, drop or rewrite the question rather than keep a broken gold label.
Keeping the eval set private and fresh
Keep the corpus and questions unpublished and refresh them on a schedule, because public multimodal benchmarks leak into training data. Researchers have documented contamination of both the text and image sides of multimodal benchmarks in model training data [8], and continuously refreshed benchmarks such as MMBench-Live exist precisely because static sets age [9]. An enterprise eval set built from non-public files avoids most of that, as long as you never paste it into third-party tools that retain prompts.
Operationally: store the set in an access-controlled bucket, log who runs it, rotate in a new quarter of questions each cycle while retiring a slice, and keep a small fixed anchor slice for trend lines. For the broader approach to held-out sets, see private evaluation sets for multimodal models and AI evaluation datasets built from real business work.
Specifying a request for multimodal RAG evaluation content
A good request describes file types, modality mix, domain and permitted use, not a particular company. Before sourcing, write down the target distribution per modality tag, the minimum share of questions whose answer is absent from the text layer, the annotation depth (page, slide or bounding box), and whether the license must cover evaluation only or also grounding in production.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Request field | Example entry |
|---|---|
| Content | Quarterly business-review decks, architecture diagrams, internal tool screenshots |
| Formats | Native PPTX and PDF plus 150 dpi renders; PNG screenshots |
| Modality mix | 30% chart, 20% diagram, 20% screenshot, 15% table image, 15% text control |
| Evidence granularity | Slide or page plus bounding box |
| Hard-question share | At least half with answer absent from text layer |
| De-identification | Names, emails, account numbers and customer logos removed in pixels |
| Permitted use | Internal evaluation; grounding to be negotiated separately |
Our multimodal data specification template gives the full field list, and the multimodal data hub maps related datasets. If you would rather source the underlying files than build them, you can describe the corpus to SourceX.
Sourcing real enterprise content for multimodal RAG evaluation
SourceX sources operational datasets, including documents and engineering records, from US companies on request; nothing is held in stock and a request does not guarantee a match. Each dataset is rights-reviewed, personal details are removed or replaced before delivery, and every release is approved by the supplying company and delivered under a license that defines records, uses, term and delivery. Describe the slides, diagrams and screenshots your evaluation needs.
Sources
- arXiv, "Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation" (2026). https://arxiv.org/pdf/2604.04936
- arXiv, "EnterpriseRAG-Bench: A RAG Benchmark for Company Internal Knowledge" (2026). https://arxiv.org/pdf/2605.05253
- arXiv, "WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation" (2025). https://arxiv.org/html/2505.08643v1
- arXiv, "DocVQA: A Dataset for VQA on Document Images" (2020). https://arxiv.org/pdf/2007.00398
- arXiv, "SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images" (2023). https://arxiv.org/pdf/2301.04883
- arXiv, "MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations" (2024). https://arxiv.org/html/2407.01523
- arXiv, "ICDAR 2019 Competition on Scene Text Visual Question Answering" (2019). https://arxiv.org/pdf/1907.00490
- arXiv, "Both Text and Images Leaked! A Systematic Analysis of Multimodal LLM Data Contamination" (2024). https://arxiv.org/abs/2411.03823v1
- arXiv, "MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models" (2026). https://arxiv.org/pdf/2607.01813
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.