Multimodal and embodied data
Robot Teleoperation Data: What to Specify When Sourcing Demonstrations
Quick answer
A usable robot teleoperation dataset is defined by its embodiment, not its hour count. Specify the robot and end effector, control mode and frequency, every synchronized observation stream, the action representation, per-episode metadata (task, language instruction, success label and who judged it), operator QA, scene diversity and the license. Open sets show the structure; commercial demonstrations must match your rig and carry clear rights for training.
By SourceX Editorial · Updated
This guide sits under our multimodal and embodied data hub and focuses on one collection method: a human driving a robot through a leader arm, VR controller or spacemouse while the robot logs what it saw and did. For the broader robotics intent, see robotics training data for embodied AI.
Start with the embodiment: robot, hand, control mode and rate
Demonstrations transfer best when the recording robot matches the robot your policy will run on, so the embodiment block is the first line of any request. Name the arm model, degrees of freedom, gripper or dexterous hand, mounting (fixed table, mobile base, rolling desk) and whether the setup is single-arm or bimanual. Multi-site collections work best when every site builds the same rig, so ask whether all episodes in a delivery came from identical hardware, camera placements and firmware.
Control mode matters as much as hardware. Leader-arm bimanual rigs map the operator's joints one to one onto the follower arms, which yields joint-position actions; VR and spacemouse rigs usually command end-effector poses through inverse kinematics. A policy trained on absolute joint targets at one rate will not consume delta end-effector commands at another without conversion, so ask for:
- Action space: joint positions, joint velocities, absolute or delta end-effector pose, and the rotation convention (quaternion order, Euler sequence, axis-angle).
- Control and logging frequency for actions and each sensor, plus how they were resampled.
- Gripper action type: binary, continuous width, or per-finger joints.
- Calibration files: camera intrinsics and extrinsics, URDF version, and the base frame used.
If your team is weighing physical data against simulation for this embodiment, real versus simulated robot data covers when the physical premium pays off.
Define every observation stream and how it is synchronized
A teleoperation episode is only as good as the alignment between what the cameras saw and what the robot did at that instant. List each stream with resolution, frame rate and encoding: wrist camera(s), one or more third-person cameras, depth if present, joint positions and velocities, gripper state, and force-torque or tactile signals where the task is contact-rich (see tactile and force-torque data).
Ask how timestamps were produced. Hardware-triggered cameras and a shared clock are stronger than software timestamps stamped on arrival, and USB camera latency can vary by tens of milliseconds. Request the per-stream timestamp arrays, not just pre-aligned frames, so you can re-align or drop episodes where drift exceeds your tolerance.
Format is part of the spec. Robotics data usually ships as RLDS (TFRecord episodes), LeRobot datasets (Parquet plus MP4), HDF5 per episode, or raw ROS bags and MCAP logs; robotics dataset formats compares them. Ask for raw logs alongside any converted format, because lossy video compression or downsampling during conversion is hard to undo.
Require episode-level metadata, including who judged success
Episode metadata is what turns hours of video into trainable, filterable demonstrations. At minimum each episode needs a task ID, a free-text language instruction, a success label, a scene ID, a pseudonymized operator ID, a rig or robot serial ID and a collection date. If you are training a vision-language-action model, the instruction quality and paraphrase coverage matter as much as the trajectories; vision-language-action training data goes deeper.
Success labels need provenance. Ask whether success was judged by the operator in the moment, by a second reviewer from video, by a scripted check, or by a learned success detector; research shows vision-language models can serve as success detectors, which makes the labeling method a variable you should record rather than assume [2]. Keep failed and aborted episodes with their labels instead of letting a supplier silently drop them, since failure, intervention and recovery data is useful for both filtering and recovery behaviors.
Illustrative example: invented to show structure; it does not describe an available dataset.
episode_id: ep_000412
task_id: drawer_open_place_mug
language_instruction: "Open the top drawer and put the blue mug inside."
robot: { arm: "7-DoF arm", end_effector: "parallel gripper", serial: rig_03 }
control: { mode: delta_ee_pose, rate_hz: 15, rotation: axis_angle }
streams:
wrist_rgb: { res: 1280x720, fps: 15, codec: h264 }
ext_rgb_1: { res: 1280x720, fps: 15, codec: h264 }
joint_state: { rate_hz: 100 }
gripper: { type: continuous_width, rate_hz: 100 }
sync: { clock: shared_ptp, max_drift_ms: 12 }
scene_id: kitchen_b_left_counter
lighting: overhead_fluorescent_plus_window
operator_id: op_7f2c # pseudonymized
operator_episode_index: 38
success: true
success_judged_by: second_reviewer_from_video
interventions: 0
duration_s: 41.6
Measure operator quality, not just operator count
Demonstration quality varies by operator, and that variance shows up directly in imitation learning. Human demonstrations are non-stationary: the same person does the same task differently across attempts, pauses, corrects and speeds up as they learn the rig, and behavior-cloning errors compound over long horizons. Treat operators as a measured input.
Ask suppliers for operator-level statistics: episodes per operator, success rate, median episode duration and its spread, intervention or reset counts, and how many hours each operator had on the rig before recorded sessions. Long sessions degrade quality, so request session start times or an episode index within the session so you can check for fatigue effects. Pause-heavy or hesitant trajectories (long idle segments, jerky corrections) are worth flagging with a simple velocity-profile filter during acceptance.
Prioritize scene and object diversity over repeated takes
For generalist policies, varied scenes, objects and lighting usually buy more than hundreds of repeats of one tabletop. Large open efforts were built on this premise: Open X-Embodiment pooled over one million trajectories from 22 embodiments to test cross-robot transfer [1]. A commercial request should state which axis of diversity it is paying for.
Write diversity into the request as countable targets: distinct scenes, distinct object instances per category, lighting conditions, camera placements, and distractor clutter levels. The opposite regime also exists: narrow, fine-grained bimanual skills on a single fixed rig can often be learned from a small number of high-quality demonstrations. Decide which regime you are in before pricing hours.
Request template for teleoperated demonstrations
The fastest way to get comparable answers from suppliers is to send the same structured request to each. Fill this in before talking to anyone; buyers who later route a request through SourceX's buyer process describe the data in this way rather than naming target businesses.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | What to state | Why it matters |
|---|---|---|
| Embodiment | Arm model, DoF, end effector, single or bimanual, fixed or mobile | Transfer degrades across embodiments |
| Teleop interface | Leader arm, VR controller, spacemouse, exoskeleton glove | Shapes action smoothness and speed |
| Action space and rate | Joint or end-effector, absolute or delta, Hz | Must match the policy head |
| Streams | Cameras with placement, depth, proprioception, force-torque | Defines model inputs |
| Sync | Clock source, max allowed drift, timestamp arrays | Misalignment corrupts labels |
| Tasks | Task list, language instruction style, paraphrases | Drives instruction following |
| Success labeling | Judge, criteria, inter-rater check | Filtering and evaluation |
| Diversity | Scenes, object instances, lighting, clutter | Generalization |
| Operator QA | Per-operator stats, training hours, session length | Demonstration consistency |
| Failures | Keep aborted and failed episodes, labeled | Recovery and filtering |
| Format | RLDS, LeRobot, HDF5, MCAP, plus raw logs | Pipeline fit |
| Privacy | Faces, voices and screens in frame; treatment method | Rights and biometric law |
| Rights | Who owns the data, allowed uses, term, delivery | Licensability |
Before payment, run the checks in acceptance checks for robot datasets: replay a sample of actions in simulation or on your rig, verify timestamp alignment, and recompute success rates on a held-out sample.
Check rights: open licenses, rig owners and people in frame
Teleoperation data has several possible rights holders, so the license question is not settled by who pressed record. Robot telemetry has no simple statutory owner; access and use are typically allocated by contract, and in the EU the Data Act, applicable since 12 September 2025, adds access rules for connected-product data as of October 2026 [5]. Expect the operator employer, the facility owner and sometimes the robot OEM to have a say; who owns robot data maps these parties.
Open datasets are a starting point, not a license to ship. Practitioner guidance comparing DROID with commissioned collection stresses checking each dataset's terms before commercial use [3], and open robot datasets and commercial licenses walks through common license patterns.
Third-person cameras often capture operators' faces and voices. Under Illinois BIPA, as amended by SB 2979, repeated collection of the same biometric identifier from the same person by the same method counts as one violation, which lowers but does not remove exposure [4]. Ask how operators consented, whether faces are blurred or cropped, and whether audio was removed. This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
License existing recordings or commission new collection
Licensing existing recordings is faster when a business already runs a matching rig; commissioning is the only option when your embodiment or task set is new. Many companies running pilots, R&D cells or deployed robots hold teleoperation logs as a by-product, but their rigs, rates and labeling rarely match yours exactly. Commissioning versus licensing robot data compares both paths, and related demonstration sources include egocentric video of skilled manual work and our physical-world data overview.
SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and manages the licensing and ongoing purchases. Nothing is held in stock and a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows after an executed agreement and supplier approval. SourceX does not train models.
Sourcing teleoperated robot demonstrations
If you need teleoperated demonstrations that open datasets do not cover, describe the embodiment, streams, tasks and rights you require. SourceX looks for US businesses that hold matching data, assesses the data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Describe the demonstration data you need.
Frequently asked questions
Is VR teleoperation data interchangeable with leader-arm data?
Not directly. VR controllers usually command end-effector poses through inverse kinematics, while leader arms map joints one to one, so action spaces, smoothness and achievable dexterity differ. Mixing them requires converting actions to a shared representation and tagging the interface per episode.
Should failed teleoperation episodes be delivered?
Yes, labeled as failures with the reason if known. They let you filter cleanly, train success detectors and study recovery, and their absence can hide an inflated success rate.
How much bimanual data do we need?
It depends on task difficulty and policy architecture. Narrow skills on one fixed rig can need little data, while generalist models rely on broad, multi-embodiment corpora [1]. Pilot with a small set before committing to volume.
Sources
- Open X-Embodiment Collaboration (arXiv), "Open X-Embodiment: Robotic Learning Datasets and RT-X Models" (2023). https://arxiv.org/abs/2310.08864v1
- arXiv, "Vision-Language Models as Success Detectors" (2023). https://arxiv.org/pdf/2303.07280
- RoboticsCenter.ai, "DROID Dataset Buying Guide". https://www.roboticscenter.ai/guides/robot-data/droid-dataset-buying-guide.html
- Illinois General Assembly, "SB 2979 (103rd General Assembly), BIPA amendment, engrossed text" (2024). https://www.ilga.gov/documents/legislation/103/SB/PDF/10300SB2979eng.pdf
- CMS, "Data Act and Physical AI: Who owns the data from connected robots?". https://cms.law/en/deu/legal-updates/data-act-and-physical-ai-who-owns-the-data-from-connected-robots
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.