Multimodal and embodied data
Writing a Multimodal Data Request: A Specification Template
Quick answer
A multimodal data request should define the model's purpose and then specify, for each modality, what counts as one record, how the modalities are paired and kept in sync, the formats and metadata, the permitted uses and consents, and the de-identification steps. It also needs a sample stage with measurable acceptance criteria. Two thresholds do most of the work: a maximum missing-modality rate and a sync tolerance. The template below can be copied into an RFP or a supplier brief.
By SourceX Editorial · Updated
A single-modality request can get by with a schema and a volume. A multimodal request cannot, because most failures sit between the streams: video without its audio track, transcripts timed to the wrong clock, robot joint states logged at 50 Hz against a 30 fps camera with no shared timestamp, or a screen recording that the meeting host licensed but the presenters did not. The general guide on how to write a data request for suppliers covers scope and commercial basics. This page adds the fields that only multimodal data needs. For the wider cluster, start at the multimodal training data hub.
Purpose and model: what the data has to teach
Start with the model behavior the data supports, because that choice sets the pairing, sync and volume requirements. A vision-language model trained on interleaved documents needs image-text adjacency, not millisecond sync. A vision-language-action policy needs frame-accurate alignment between camera frames, proprioception and actions. A multimodal RAG evaluation set needs held-out question-answer pairs grounded in slides and screenshots.
State three things up front: the task (pre-training, supervised fine-tuning, evaluation or retrieval), the target model's input contract (for example "two RGB views at 224x224 plus a language instruction, predicting 7-DoF end-effector deltas"), and the deployment setting. Suppliers use this to say whether their recordings fit before anyone discusses formats. If the request is for held-out testing, read private multimodal evaluation sets first, since contamination rules change the request.
Modalities and pairing: defining one record
Define the record unit explicitly, because "a video with transcript" can mean a whole meeting, a speaker turn or a five-second clip. For each modality, list whether it is required or optional, its role (input, target or context), and the join key that ties it to the others. Typical units are an episode (robotics), a session (meetings, support calls), a ticket with attachments, or a document page.
Pairing has two failure modes worth naming in the request. Loose pairing means items share a parent ID but nothing proves they describe the same moment, such as a screenshot attached to a ticket three replies after the text it illustrates. Orphaned modalities are records missing a required stream. Ask suppliers to report the share of records with each modality present, and set a maximum missing-modality rate per required stream rather than one blended number. For image-text pairs, the checks in image-text alignment quality checks help write measurable pairing criteria.
Sync and calibration: tolerances you can test
Sync requirements should be a number with a unit, measured against a named reference clock. Write "audio-video offset no greater than 40 ms measured against the camera's presentation timestamps" rather than "synchronized." For robotics and sensor fusion, ask whether streams share a hardware clock or were aligned after the fact, which timestamp each sample carries (capture, receive or log), and how dropped frames are flagged.
Calibration belongs in the same section. For multi-camera or camera-LiDAR rigs, request intrinsics, extrinsics and the date each calibration was taken, plus a flag when a rig changed mid-collection. For robots, request the URDF or kinematic description, action space definition, control frequency and gripper encoding. The deeper buyer checks are in time synchronization in multimodal datasets and camera, LiDAR and radar fusion data.
Formats and packaging: say what your loader reads
Name the container and per-modality encodings your pipeline already reads, and accept conversion only where you can verify it. Common choices are MP4 (H.264 or H.265) or MKV for video, WAV or FLAC for audio at a stated sample rate, JSON Lines for text and labels, Parquet for tabular telemetry, and WebVTT or JSON word-level timings for transcripts. For large training sets, WebDataset stores samples in numbered tar shards, and files sharing a basename are grouped as one sample, which makes pairing visible in the file layout [3].
Robot data is often logged as ROS bags or MCAP and later converted to an episode format such as RLDS or LeRobot. The Open X-Embodiment effort pooled datasets from many robot types only after converting them into a common episode format [4]. If you need conversion, specify which side performs it and require the raw logs for a sample so you can check what the conversion dropped. See what file formats AI buyers accept for a broader format list, and the technical delivery specification template for transfer, checksums and manifests.
Metadata and documentation: fields per record and per dataset
Ask for metadata at two levels. Dataset-level documentation should follow an established structure: Datasheets for Datasets covers motivation, composition, collection process and recommended uses [5], and Croissant-RAI expresses responsible-AI metadata in a machine-readable form [6]. Record-level metadata should let you filter, slice and audit without opening the media.
Useful record-level fields include capture device and firmware, capture date (or a coarsened period), duration, frame rate and resolution, language and locale, environment or site type (not the site's identity), task or scenario label, collection mode (operational log, teleoperation, scripted recording), consent basis code, de-identification method code and a quality flag. For teams subject to EU AI Act Article 10 for high-risk systems, record-level provenance and documented data preparation make the required governance easier to show [10]. As of October 2026, Article 10 is amended by Regulation (EU) 2026/1744 [11], which also moved the high-risk application dates.
Rights and consents: per modality, not per dataset
Ask for a rights statement per modality, because one session often has several rightsholders. In a recorded meeting, the company may own the recording while external participants consented only to internal use, and a shared slide may contain third-party material. A robot episode may involve an operator's logs, an OEM's telemetry format and a facility's camera footage. The cluster guides on licensing records from several rightsholders and robot data ownership cover these splits.
State permitted uses per modality: training, evaluation, retrieval indexing, or excluded. Faces and voices need particular care. Illinois BIPA regulates collection, retention and disclosure of biometric identifiers such as face geometry and voiceprints and requires written release [7]. Ask suppliers whether any stream was used to derive biometric templates and how participant notice was given.
De-identification across streams
De-identification has to cover each stream and the links between them, because a name scrubbed from the transcript can still be spoken in the audio, visible on a name badge in frame, or shown in a screen-share title bar. Ask for the method per modality: face and license-plate blurring, voice masking or audio redaction of spoken identifiers, OCR-based redaction of on-screen text, and stripping of EXIF, GPS and device serial metadata.
Health data adds formal rules. HIPAA's de-identification standard lists full-face photographic images and biometric identifiers among the identifiers that must be handled [8], and medical imaging needs DICOM header handling, where the standard itself notes its confidentiality profiles do not guarantee removal of all identifying information [9]. Request the redaction log and a reviewed sample instead of a statement that data is "anonymized." Details are in de-identifying multimodal records.
Sample stage and acceptance criteria
Plan a sample stage with at least one revision cycle before full delivery; vendor practice commonly runs inquiry, sample and iteration before licensing the full set [1]. Samples are often under evaluation-only terms that bar other use and redistribution, as in NVIDIA's sample data license [2], so confirm your evaluation fits those terms before running training experiments. A sample manifest helps both sides agree what the sample contains.
Write acceptance criteria before you see the sample, and measure the sample with the same scripts you will run on full delivery. Each criterion needs a metric, a threshold and a remedy, such as replacing failed records. Once the specification is drafted, you can submit it as a buyer request to SourceX to look for US companies that hold matching data.
The template
Illustrative example: invented to show structure; it does not describe an available dataset.
request_id: MM-REQ-001
purpose:
task: fine-tuning # pre-training | fine-tuning | evaluation | retrieval
model_input: "2 RGB views 224x224 + instruction text -> 7-DoF action"
deployment: "warehouse picking, mixed SKUs"
record_unit: episode # episode | session | ticket | page | clip
modalities:
- name: wrist_camera
required: true
role: input
format: "MP4 H.264, 30 fps, 640x480"
- name: joint_states
required: true
role: input
format: "Parquet, 50 Hz, radians"
- name: actions
required: true
role: target
format: "Parquet, 10 Hz, end-effector deltas + gripper 0/1"
- name: instruction_text
required: true
role: context
format: "JSONL, English"
- name: audio
required: false
role: excluded # not requested; strip if present
pairing:
join_key: episode_id
max_missing_rate: {wrist_camera: 0.01, joint_states: 0.00, actions: 0.00, instruction_text: 0.02}
sync:
reference_clock: robot controller monotonic clock
tolerance_ms: 20
dropped_frames_flagged: true
calibration: [camera_intrinsics, camera_extrinsics, urdf, calibration_date]
volume_and_diversity:
min_episodes: "TBD after sample"
coverage_slices: [object_category, lighting, operator_id_hashed, success_flag]
max_share_any_single_site: 0.25
packaging: "WebDataset tar shards; raw MCAP for sample only"
metadata_per_record: [device, firmware, capture_period, collection_mode,
consent_code, deid_method_code, quality_flag]
documentation: [datasheet, croissant_rai_json, conversion_notes]
rights:
per_modality_uses:
wrist_camera: [training, evaluation]
instruction_text: [training, evaluation]
biometric_templates_derived: false
deidentification:
faces: blur, verified on reviewed sample
on_screen_text: OCR redaction
file_metadata: strip EXIF/GPS/serials
acceptance:
- {metric: missing_modality_rate, threshold: "per pairing.max_missing_rate", remedy: replace}
- {metric: av_or_sensor_offset_p99_ms, threshold: 20, remedy: replace}
- {metric: decode_failure_rate, threshold: 0.001, remedy: replace}
- {metric: residual_identifier_rate_in_review, threshold: "0 in reviewed sample", remedy: re-process}
- {metric: duplicate_episode_rate, threshold: 0.005, remedy: remove}
sample:
size: "50 episodes across all coverage slices"
terms: evaluation-only, no redistribution
iterations: 1-2
Adapt the modality list to the domain. A meeting dataset swaps joint states for speaker-diarized transcripts and screen-share frames (multimodal meeting recordings); a support dataset swaps them for ticket text and attached screenshots (support tickets with screenshots).
Request multimodal data through SourceX
SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents and new recordings of hands-on work; nothing is held in stock, and a request does not guarantee a match. Each dataset is rights-reviewed, has personal details removed or replaced before delivery, and is delivered under a license defining records, uses, term and delivery, after the supplying company approves the release. To start, describe the multimodal data you need on the buyers page.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Frequently asked questions
How strict should the sync tolerance be?
Set it from the model's input rate, not from what sounds rigorous. A policy acting at 10 Hz rarely needs sub-millisecond alignment, while lip-sync or audio-visual speech models need offsets well under one video frame. Measure tolerance as a high percentile (for example p99) rather than a mean, because a few badly aligned episodes can hide inside a good average.
Should one request cover several modalities' rights at once?
Yes, but keep the rights table per modality. The license can be one document, yet the permitted uses, consent basis and de-identification method usually differ between video, audio, text and telemetry, and reviewers need to see each separately.
What if a supplier holds most but not all required modalities?
Decide in advance which modalities are hard requirements. Marking a stream as optional, with its own missing-rate threshold, lets suppliers propose partial matches you can assess rather than declining outright.
Sources
- Pocstock, "The dataset licensing process: from inquiry to delivery". https://support.pocstock.com/en/articles/14772883-the-dataset-licensing-process-from-inquiry-to-delivery
- NVIDIA, "NVIDIA Sample Data License for Evaluation" (2026). https://developer.download.nvidia.com/licenses/nvidia-sample-data-license-for-evaluation-2026.01.19.pdf
- WebDataset project, "webdataset (GitHub repository)". https://github.com/webdataset/webdataset
- Open X-Embodiment Collaboration, "Open X-Embodiment: Robotic Learning Datasets and RT-X Models" (2023). https://arxiv.org/html/2310.08864v8
- Gebru et al., "Datasheets for Datasets" (2021). https://arxiv.org/pdf/1803.09010
- Jain et al., "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
- Illinois General Assembly, "Biometric Information Privacy Act, 740 ILCS 14/15". http://www.ilga.gov/legislation/ilcs/fulltext.asp?DocName=074000140K15
- eCFR, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- NEMA / DICOM Standards Committee, "DICOM PS3.15 Annex E: Attribute Confidentiality Profiles" (2026). https://dicom.nema.org/medical/dicom/current/output/chtml/part15/chapter_E.html
- European Commission, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- European Parliament and Council of the European Union (EUR-Lex), "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.