Multimodal and embodied data
Private Evaluation Sets for Multimodal Models: Sourcing Held-Out Image, Audio and Video Tasks
Quick answer
A useful multimodal LLM evaluation dataset is one your model has never seen: real images, audio and video from a working domain, paired with gold answers that come from recorded outcomes rather than annotator guesses, and licensed under terms that keep it out of training. Public benchmarks leak into training corpora, so release gating needs a private set. Source it from organizations that hold the media, run text and media contamination checks before acceptance, write eval-only use into the license, and plan a refresh before the set goes stale.
By SourceX Editorial · Updated
Why public multimodal benchmarks stop measuring what you think
Public multimodal benchmarks lose validity because both their text and their media end up in training corpora. Leakage can enter through the base language model, which may have memorized question and answer text, or through the multimodal training stage, which may have seen the images, clips or captions themselves. A tell-tale symptom is a model that answers "visual" questions correctly with the image removed.
Because most VLM training mixtures are proprietary, you usually cannot audit leakage from the outside, and inflated scores can go unnoticed. Market practitioners make the same point about text benchmarks: once a test set is public it drifts into crawled data, and a private set only protects you while it stays unpublished [1]. For background on when held-out data is worth the cost, see private evaluation sets vs public benchmarks and the glossary entry on held-out data.
There is a second, quieter problem. Label errors in public test sets, including audio and vision sets, were estimated at an average of at least 3.3% in one audit of ten widely used benchmarks [4]. When model deltas between releases are a point or two, a noisy gold set can reverse a ship decision.
What a private multimodal eval set should contain
A private multimodal eval set is a versioned collection of tasks, each with its media, a prompt, a gold answer, the evidence for that answer and the rights metadata that governs use. The strongest sets come from operational records where the "answer" was settled by what actually happened: the part the technician replaced, the defect code the inspector logged, the resolution a support agent applied after reading a screenshot.
Useful task families by modality:
- Image and document: photos of equipment with the logged fault, scanned forms with the keyed values, screenshots from support tickets with attachments paired with the resolution.
- Audio: call or field recordings with the outcome recorded in a CRM or work order, plus speaker-turn and timestamp metadata.
- Video: line or procedure video with quality outcomes from an MES or QA log, or meeting recordings with decisions captured in minutes.
- Cross-modal: tasks that require two inputs, such as a slide image plus the spoken explanation, where neither alone answers the question. A useful filter: drop any item a text-only model can answer with the media removed.
- Agentic: where you test computer-use agents, execution-based checking against a real environment, as the original OSWorld release does across 369 tasks, is stronger than string matching [5].
For retrieval-heavy products, the sibling guide on evaluating multimodal RAG on enterprise content covers slides, diagrams and screenshots as a corpus rather than as single tasks.
Gold answers from outcomes, not opinions
Gold answers are most defensible when they are taken from a system of record and verified against the media, not written fresh by an annotator. Ask the supplier which field the answer came from (for example work_order.resolution_code or qa_inspection.defect_class), when it was written relative to the media, and whether anyone could have seen the outcome while capturing the media.
Then plan for disagreement. Have two domain reviewers check a sample of answers against the media; route conflicts to an adjudicator; and record answer_confidence and adjudicated flags per item. Where outcomes are ambiguous, keep the item but mark it as a soft-scored or rubric item rather than forcing a single label. Our guide to sourcing domain-expert raters covers reviewer qualification.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"task_id": "mmeval-v3-000412",
"modalities": ["image", "audio"],
"media": [
{"uri": "shard-0007.tar#000412.jpg", "sha256": "…", "phash": "c3a1f0e…", "captured_at": "2025-03-14T10:22:05Z"},
{"uri": "shard-0007.tar#000412.flac", "sha256": "…", "duration_s": 41.8, "sample_rate_hz": 16000}
],
"prompt": "From the photo and the technician's voice note, which component failed?",
"gold_answer": "Condenser fan motor",
"answer_type": "exact_match_with_synonyms",
"answer_source_field": "work_order.resolution_code",
"answer_written_at": "2025-03-14T15:40:00Z",
"adjudicated": true,
"text_only_baseline_correct": false,
"deidentification": {"faces": "blurred", "voice": "unchanged_no_names_spoken", "exif": "stripped"},
"license_use": "evaluation_only",
"split": "gating",
"set_version": "3.0"
}
Contamination checks for images, audio and video
Media contamination checks look for your eval items, or near copies of them, anywhere a training crawler could have found them. Text-only n-gram checks miss reused images and clips, so a multimodal acceptance test needs per-modality fingerprints plus behavioral probes.
| Check | What it catches | How to run it |
|---|---|---|
| Perceptual hash (pHash, dHash) against public crawl snapshots and your own training manifests | Resized, recompressed or lightly cropped images | Hash every frame or image; flag Hamming distance below a set threshold |
| Reverse image search on a sample | Images already published on the web, such as product pages or forums | Manual or API lookups; any hit removes the item |
| Audio fingerprinting | Recordings reused in podcasts, training corpora or demos | Chromaprint-style fingerprints against known corpora |
| Video keyframe hashing | Clips posted publicly or reused in marketing | Hash keyframes at a fixed interval; compare to crawl and training indexes |
| Text-only and blind baselines | Leaked questions or answers | Run the prompt with media removed; correct answers suggest the item leaked or never needed the media |
| Perturbation probes | Memorized items in a model you cannot audit | Shuffle multiple-choice options, mask caption words for slot-guessing, or perturb the media semantically and watch for answers that do not change |
Run these before you accept delivery, and again on any model whose training data you did not control. When you describe the evaluation data you need, list the checks you will run so suppliers know the acceptance bar. SourceX's guide to contamination checks for licensed eval data covers the text side; the broader quality and contamination hub covers training sets.
Contract terms that keep a held-out set held out
A held-out set stays held out only if the license and your own controls stop it from reaching training pipelines, yours or anyone else's. Ask for terms that cover:
- Permitted use: evaluation, regression testing and release gating; explicitly no training, fine-tuning, distillation or prompt-tuning on items or derived outputs.
- Exclusivity of the items: whether the supplier may license the same media to other model developers, which would put your set inside a competitor's training data.
- Curator independence: Bansal and Maini warn that evaluators who also sell training data face conflicts of interest and opacity [2]; require disclosure of any training-data relationships and a right to audit how items were selected.
- Access controls: named users, logged access, storage in a separate bucket from training data, and no copies in shared notebooks or vendor eval harnesses that log inputs.
- Publication: whether you may report scores, show example items in a model card, or never disclose items.
- End of term: deletion or return of media and any cached embeddings, with a written certificate.
The SourceX answer on licensing data for evaluation only and the model evaluation teams guide go further on internal controls. If you work with third-party evaluators, see data rights for independent AI evaluation organizations.
Privacy and rights in eval media
Eval media carries the same privacy load as training media, and a held-out set is often reviewed by more people. Faces, voices, on-screen names, license plates and EXIF GPS tags all need handling, and blurring one modality while leaving a spoken name in the audio is a common failure. The guide to de-identifying multimodal records covers faces, voices, screens and metadata together.
Check that de-identification does not destroy the signal the task needs: a blurred gauge face or a pitch-shifted voice can make the gold answer unrecoverable. Record the method per item, as in the deidentification field above. Composite records drawn from several rightsholders, such as a video plus a third party's slide deck, need each component cleared; see licensing multimodal records from several rightsholders.
Sizing, splits and refresh cadence
Size the set for the decisions it gates, then split it so one leak does not burn everything. A common pattern is a small dev split engineers can inspect, a larger gating split only the eval harness reads, and a reserve split held back for the next version. LiveBench's authors argue that contamination can make a benchmark obsolete quickly and use regular refreshes as the remedy [3]; apply the same thinking to private sets by versioning (set_version), retiring any item that appears in a leak or a vendor log, and adding fresh items from newer operational records each cycle.
For delivery, ask for media in WebDataset tar shards with a manifest carrying SHA-256 hashes, and check time synchronization for any task that pairs audio, video and logs.
Buyer checklist for a private multimodal eval set
Illustrative example: invented to show structure; it does not describe an available dataset.
| Area | Ask the supplier | Accept when |
|---|---|---|
| Provenance | Which system produced the media and the answer field? | Both are named; answer timestamp follows media capture |
| Novelty | Has any item been published, shared or licensed before? | Written confirmation plus clean hash and reverse-search results |
| Gold quality | How were answers verified? | Double review on a sample, adjudication log provided |
| Vision dependence | Do text-only baselines answer items? | Flagged items removed or marked |
| Privacy | Which de-identification method, per modality? | Method recorded per item; sample checked |
| Rights | Who owns each component and what consents cover it? | Every component cleared for evaluation use |
| Use terms | Training, distillation, publication, deletion | Eval-only use written into the license |
| Refresh | Can new items be added on a schedule? | Versioned additions with the same checks |
Sourcing private multimodal evaluation data with SourceX
SourceX sources operational datasets, including support histories, engineering records, documents and new recordings of hands-on work, from US companies on request; a request does not guarantee a match, and nothing is contracted until the supplying company agrees. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery. For background, see the multimodal data hub, the AI evaluation data overview and evaluation sets built from real business work, then describe the multimodal evaluation data you need.
Frequently asked questions
How large should a private multimodal eval set be?
Size it from the smallest score difference you need to detect between releases and the number of slices you report (modality, domain, difficulty). Small, carefully adjudicated sets often beat larger noisy ones, given that public test sets carry measurable label error [4].
Can we reuse items from our training data vendors as eval items?
Only if you can prove the items never entered any training mixture, including the vendor's other customers. In practice, sourcing eval media from a separate supplier and contract reduces that risk.
Should we publish part of the set?
Publishing a small sample helps credibility, but treat published items as burned for gating and move them to a public demo split.
Sources
- DataForce, "Disadvantages of standard LLM benchmarks". https://www.dataforce.ai/blog/disadvantages-standard-llm-benchmarks
- arXiv (Bansal and Maini; ICLR 2025), "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
- arXiv (ICLR 2025), "LiveBench: A Challenging, Contamination-Free LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
- arXiv (Northcutt, Athalye, Mueller; NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- arXiv (Xie et al.; NeurIPS 2024), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.