Multimodal and embodied data
Vision-Language-Action Training Data: Pairing Robot Trajectories with Language
Quick answer
A vision language action dataset pairs synchronized camera frames and robot actions with natural-language instructions, so a policy learns to map "what it sees" plus "what it was told" to motor commands. For VLA pre-training or fine-tuning, specify four layers: the trajectory itself, a task-level instruction, step-level narration with failure descriptions, and embodiment metadata (robot type, action space, camera placement) that lets you pool data. Then hold out unseen instructions and objects for evaluation before you buy anything.
By SourceX Editorial · Updated
What makes trajectory data "VLA-ready"
VLA-ready data is trajectory data in which every episode carries language that actually conditions the actions, not a filename or a task ID. A common VLA recipe co-fine-tunes a vision-language model on robot trajectories alongside web vision-language tasks, with actions written as text tokens in the same training stream as words. In that setup the instruction is a direct conditioning input alongside the image, so a vague or templated instruction limits what the policy can learn to follow.
Generic video or image-caption corpora rarely work as substitutes because they lack action grounding: the commanded joint or end-effector deltas, gripper state and outcome of each step [8]. Pooled open collections such as Open X-Embodiment, which gathers over one million trajectories from 22 embodiments [1], and DROID, a teleoperated in-the-wild manipulation set [2], are the usual starting point. Commercial buyers typically want what those sets lack: real tasks in real facilities, task vocabulary from skilled work, and license terms that cover commercial training. See open robot learning datasets and their commercial licenses before you assume a public set is usable.
The four language layers to specify
Specify language as separate layers with their own fields, because each layer trains a different capability. Collapsing them into one free-text "description" column is a common reason a purchased set cannot be reused for a second model.
- Task instruction. One imperative sentence per episode ("place the torque wrench in the second drawer"). This is what an instruction-following policy is conditioned on at inference.
- Step-level narration. Timestamped sub-goals ("grasp handle", "pull drawer open") aligned to frame or step indices. AndroidControl, a UI-agent dataset, pairs each episode with a high-level goal and per-step low-level instructions, and that two-level design is a useful template for manipulation too [3].
- Failure and recovery descriptions. Text that says what went wrong ("grasp slipped; re-approached from the left"). These pair well with robot failure, intervention and recovery data.
- Paraphrases. Several phrasings of the same instruction, written independently, to raise instruction diversity without collecting more motion. Mark which paraphrase was shown to the operator, if any.
Step boundaries should come from an explicit segmentation spec. The same issues that arise in temporal action segmentation labels (boundary tolerance, label ontology, overlapping actions) apply to narration.
Who writes the language changes quality and rights
The author of each language string is a provenance field, not a detail. An operator who speaks the instruction before teleoperating produces language that matches intent but may be terse; an annotator writing after the fact produces cleaner text that may describe what happened rather than what was asked; a model-generated caption scales cheaply but can hallucinate objects and may carry the generating model's output terms.
Record language_author (operator, annotator, model), the annotation tool and guideline version, and whether a human reviewed model output. Croissant-RAI provides a machine-readable place for labeling and life-cycle documentation of this kind [6]. For annotator agreements and AI-assistance disclosure, see provenance for human-annotated data; for who owns the recordings themselves, see who owns robot data.
Embodiment metadata that lets you pool data
Embodiment metadata is what makes trajectories from different robots, grippers and camera rigs trainable together. Open X-Embodiment pooled data across 22 embodiments [1], and every cross-embodiment effort depends on knowing exactly what each action vector means.
Ask for, at minimum: robot make and model, arm DOF, gripper type, action space (joint positions, joint velocities, end-effector delta pose, absolute pose), action frame and units, control frequency, camera count with intrinsics and extrinsics (wrist, third-person, head), image resolution and frame rate, and the timestamp source used to synchronize them. Proprioception and any tactile or force-torque channels need the same treatment. Human-only footage without robot actions is a different product; see human demonstration video for robot learning.
Delivery formats and record structure
Request delivery in a sequential format that preserves episode order, because shuffled frames destroy the temporal signal. RLDS represents a dataset as episodes of ordered steps, where each step holds observation, action, reward and discount, with is_first and is_last boundary flags and user-defined metadata at both episode and step level. Some catalog listings for VLA benchmark data offer the same set in RLDS, LeRobot and HDF5 variants [7], so name the format your training stack loads.
Describe the dataset in Croissant JSON-LD so loaders and reviewers can read field types without opening files [5]. Keep language in step metadata rather than burned into video, and keep raw frames separate from derived tokens so you can re-tokenize later.
Illustrative example: invented to show structure; it does not describe an available dataset.
episode:
episode_id: ep_000412
task_instruction: "Put the 10 mm socket back in the labeled tray"
instruction_paraphrases:
- "Return the 10 mm socket to its tray slot"
- "Store the small socket in the tray marked 10"
language_author: operator_spoken_then_transcribed
annotation_guideline_version: v3.2
sop_reference: "WI-114 tool return procedure (supplier SOP)"
success: false
failure_note: "Socket dropped at handoff; operator re-grasped and completed"
embodiment:
robot: 7-DOF arm, parallel-jaw gripper
action_space: end_effector_delta_pose_6d + gripper_binary
action_frame: robot_base
control_hz: 15
cameras:
- { name: wrist, res: 640x480, fps: 15, extrinsics: calib_0412.json }
- { name: exterior_left, res: 1280x720, fps: 15, extrinsics: calib_0412.json }
split: train # train | heldout_instruction | heldout_object | heldout_scene
steps:
- { step: 0, is_first: true, substep: "approach socket", t_ms: 0 }
- { step: 37, substep: "grasp socket", t_ms: 2466 }
- { step: 61, substep: "recover dropped socket", t_ms: 4066 }
- { step: 118, is_last: true, substep: "release into slot", t_ms: 7866 }
Real task vocabulary from written procedures
Written work procedures can supply realistic instruction language for skilled tasks, which is hard to invent. Standard operating procedures and work instructions name the tools, parts, tolerances and sequences that operators actually use, so instructions drawn from them test whether a policy understands domain vocabulary rather than tabletop toy phrasing.
If you pursue this, ask whether the supplier can link each episode to the procedure step it performed (as sop_reference above) and whether those documents can be licensed alongside the recordings. Related resources: SOP and knowledge base datasets, assembly demonstrations with work instructions and CAD and egocentric video of skilled manual work.
Holding out instructions and objects for evaluation
Build evaluation splits before training data is delivered, and split by instruction, object and scene rather than by random episode. Random splits leak near-duplicate episodes from the same session into test, which inflates instruction-following scores.
- Unseen instructions: paraphrases and new task compositions never shown in training.
- Unseen objects: object instances or categories absent from training episodes.
- Unseen scenes or embodiments: different site, lighting or robot.
AndroidControl reports that out-of-domain generalization is much harder to buy with data scale than in-domain performance [3], a pattern worth assuming for physical tasks until your own results show otherwise. Success labels can be human-judged or scored by a vision-language success detector [4]; record which. For sealed test sets, see private evaluation sets for multimodal models, and for how much data a fine-tune needs, see VLM fine-tuning data volume.
Buyer checklist for a VLA data request
Use this checklist to turn a research need into a request a supplier can answer. It also forms the basis for acceptance checks before you pay.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | What to specify | Failure mode if omitted |
|---|---|---|
| Task scope | Task families, environments, success definition | Data skews to easy pick-and-place |
| Language layers | Instruction, step narration, failure notes, paraphrase count | Single templated caption per episode |
| Language author | Operator, annotator or model; guideline version | Unknown quality; unclear model-output terms |
| Embodiment | Robot, action space, frame, Hz, camera calibration | Actions cannot be pooled or normalized |
| Format | RLDS, LeRobot or HDF5; Croissant metadata | Weeks of conversion; lost step order |
| Splits | Held-out instructions, objects, scenes | Inflated evaluation from leakage |
| De-identification | Faces, voices, screens, badges in frame | Personal data in training frames |
| Rights | Recording ownership, facility consent, allowed uses | Data cannot be used commercially |
Faces, voices and visible screens in facility footage need handling before delivery; see de-identifying multimodal records. Whether to commission new collection or license recordings that already exist is covered in commissioning versus licensing robot data and robot training data cost drivers.
Where SourceX fits for VLA trajectory data
SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and manages licensing and ongoing purchases; it holds no stock, and a request does not guarantee a match. Buyers describe the data they need, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, with the method recorded and a sample checked, though no method is perfect. Broader context is on the robotics and embodied AI use-case page and the multimodal and embodied data hub; to start a request, go to the SourceX buyer page.
Request language-annotated robot trajectories
Describe the tasks, language layers, embodiment and evaluation splits you need, and SourceX will look for US businesses that hold matching recordings or can create them. Nothing is contracted until a supplier agrees, and pricing and allowed uses are set per deal in a license. Describe the VLA trajectories you need.
Sources
- Open X-Embodiment Collaboration, arXiv, "Open X-Embodiment: Robotic Learning Datasets and RT-X Models" (2023). https://arxiv.org/abs/2310.08864v1
- Khazatsky et al., arXiv, "DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset" (2024). https://arxiv.org/abs/2403.12945v1
- Li et al. (Google DeepMind), arXiv, "On the Effects of Data Scale on UI Control Agents" (2024). https://arxiv.org/abs/2406.03679
- Du et al. (DeepMind), arXiv, "Vision-Language Models as Success Detectors" (2023). https://arxiv.org/pdf/2303.07280
- Akhtar et al. (MLCommons), arXiv, "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- Jain et al. (MLCommons), arXiv, "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
- Claru (dataset catalog), "VLA-Arena L1 L (RLDS)". https://claru.ai/datasets/vla-arena-vla-arena-l1-l-rlds
- Digital Divide Data, "Build Better Vision-Language-Action Models With Better Data". https://www.digitaldividedata.com/autonomy/vision-language-action-model
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.