Multimodal and embodied data
Real-World vs Simulated Robot Data: When to Pay for Physical Demonstrations
Quick answer
Pay for real-world robot demonstrations where simulation is structurally weak: contact-rich manipulation, deformable objects, real sensor noise and lighting, and work around people. Use simulation for scale, randomization and coverage of rare states. Most programs end up co-training on both, but real data has two jobs simulation cannot do: defining the task distribution your simulated scenes must match, and serving as the held-out evaluation set that decides whether a policy ships.
By SourceX Editorial · Updated
Why the budget question is about coverage, not volume
The right split depends on which parts of your task distribution simulation reproduces faithfully, not on a fixed ratio. A pick-and-place policy for rigid boxes on a conveyor can lean heavily on simulated rollouts, while a policy that folds garments or seats connectors may need most of its signal from physical demonstrations.
Large real-robot datasets exist because simulation alone did not close the gap. Open X-Embodiment pooled demonstrations from many robot embodiments and labs into a standardized format to train cross-embodiment policies [1], and DROID collected teleoperated manipulation in diverse in-the-wild scenes specifically to improve robustness to new environments [2]. A recent survey frames embodied data as a layered pyramid in which abundant but indirect sources (web and human video, simulation) sit below scarce, high-fidelity real robot data, and most systems combine layers rather than choosing one [3].
The cost asymmetry is real. In the adjacent domain of digital agents, synthetic demonstrations were reported at a small fraction of the cost of human-annotated ones [5], and research on learning from human video exists largely because robot data is scarce and expensive [4]. That asymmetry is why the question is where each real demonstration buys the most, not whether to buy any.
Where simulation is weakest for robot learning
Simulation tends to break down on physics and perception details that are expensive to model and easy for a policy to exploit. Treat the list below as working hypotheses to test against your own task, not settled limits.
- Contact dynamics. Friction, compliance, stiction and impact are approximated by solvers such as those in MuJoCo, PhysX or Bullet. Insertion, screwing, wiping and in-hand regrasping are where policies trained on approximate contact most often fail on hardware. Force-torque and tactile streams are the real signal here; see tactile and force-torque data for contact-rich learning.
- Deformables. Cloth, cables, bags, food and foam need cloth or finite-element models that are slow and hard to tune. Their state is also hard to label, so even good simulated rollouts carry weak supervision.
- Lighting and sensor noise. Real RGB-D has depth holes on shiny and transparent surfaces, motion blur, rolling-shutter artifacts, auto-exposure swings and lens dirt. Domain randomization widens the visual distribution but rarely reproduces the specific failure modes of your camera stack.
- Actuation and timing. Joint backlash, controller latency, dropped frames and clock skew between camera and proprioception are often absent from simulation and present in every real log.
- Human co-workers. People walking through a cell, handing over parts or blocking a camera are hard to simulate credibly, and those interactions carry safety weight.
- Long-tail failures. Jams, slips, dropped parts and recoveries are rare in scripted simulation unless someone already knows to inject them. Real failure, intervention and recovery data shows which ones actually occur.
What real-world data must cover in a co-training mix
Real data should be concentrated on the states simulation gets wrong and on the distribution you will be evaluated against. In practice that means four uses, ranked here by how hard they are to replace.
- Held-out evaluation. Even if training is mostly simulated, the go or no-go decision should come from real episodes the model never trained on, recorded on the target embodiment and in target-like scenes. This is the least substitutable spend.
- Task distribution definition. Real recordings tell you object sets, pose ranges, clutter levels, lighting conditions and operator strategies. Simulation scenes should be built and randomized to match these statistics, not invented from scratch.
- Contact-rich and deformable segments. Fund physical demonstrations for the sub-skills listed above, with synchronized force-torque and gripper state, rather than for the whole task.
- Calibration and system identification. Short real episodes used to fit friction, mass and latency parameters improve every simulated rollout that follows.
Real data does not need to come only from paid teleoperation. Operational logs from deployed robot fleets can supply realistic state distributions and failures, and commissioned collection is covered in commissioning robot data collection vs licensing existing recordings. For demonstrations specifically, see what to specify when sourcing teleoperation data.
A budget allocation worksheet for real vs simulated data
A worksheet that scores each sub-skill on sim fidelity and deployment risk turns the split into an explicit decision. Score fidelity from evidence, such as a sim-trained checkpoint's success rate on a small real probe set, rather than from intuition.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Sub-skill | Sim fidelity (evidence) | Deployment risk | Real data role | Real share of budget | Sim role |
|---|---|---|---|---|---|
| Rigid tote pick, known SKUs | High: sim policy holds up on real probe set | Low | Eval set plus small calibration set | Low | Bulk training, pose and lighting randomization |
| Bagged and deformable items | Low: grasp success drops on hardware | Medium | Training demos with gripper force and wrist camera | High | Pretraining only |
| Connector insertion | Low: contact model mismatch | High | Demos with 6-axis force-torque at 500 Hz or above | High | Approach phase only |
| Human handover zone | Unknown: no credible sim of people | High | Eval and training episodes with people present, consented | Medium | Avoid; do not train on simulated humans alone |
| Recovery from jams | Low: failures not scripted | Medium | Fleet intervention logs plus targeted demos | Medium | Inject observed failure types into scenes |
Re-score after each training round. If a sim-heavy sub-skill degrades on the real probe set, move budget back to physical data for that sub-skill only.
How much real robot data is enough
There is no general number; the amount is whatever lets your real evaluation set produce stable estimates and closes the measured gap on low-fidelity sub-skills. Size the evaluation set first, because it fixes the floor on real spend regardless of how training goes.
For evaluation, aim for enough episodes per condition (object class, scene type, lighting band) that a few percentage points of success-rate change is distinguishable from noise, and freeze the set before training starts. For training, run small ablations: train with 0, 10 and 30 percent real episodes in the co-training mix, measure real-probe success, and stop buying more real data for a sub-skill when the curve flattens. Those ablations cost less than a blind commitment to a large collection.
Watch for two traps. Pooling data across embodiments can help [1], but real data collected on a different embodiment, gripper or camera placement often transfers less than its volume suggests. And a real dataset collected in a few tidy lab scenes can underperform a smaller set spread across many scenes, which is the reason in-the-wild collections like DROID emphasize scene diversity [2].
Rights and provenance checks for both data types
Simulated data is not automatically free of rights questions, and real data carries its own consent and ownership questions. Check both before training.
For synthetic scenes, inventory every input: 3D object and environment assets, textures, HDRI lighting maps, motion-capture clips and any real recordings used to seed or tune the generator. Third-party asset marketplaces often license models for rendering or game use with terms that do not clearly cover ML training or redistribution of derived datasets. Practitioners note that synthetic data generated from licensed or proprietary inputs raises ownership and provenance questions in diligence [6], and audits of public datasets found license metadata frequently missing or wrong [7]. Do not assume an open robot dataset is cleared for commercial use; see checking open robot dataset licenses.
For real recordings, confirm who owns the data among the robot operator, OEM and facility owner (see robot data ownership in licensing), whether workers or bystanders appear on camera and consented, and whether faces, badges and screens need masking. Ask for documentation in the spirit of Data Cards: sources, collection method, intended use and known factors affecting performance [8].
Procurement checklist for physical demonstrations
Write the real-data request around the gaps your worksheet identified, so suppliers quote against coverage, not raw hours.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Embodiment: arm model, gripper type, degrees of freedom, control mode (joint position, end-effector delta, impedance) and control rate.
- Sensors: camera count and placement (wrist, overhead, side), resolution and frame rate, depth sensor type, force-torque and tactile channels with sample rates.
- Synchronization: a shared clock across streams, stated tolerance, and per-frame timestamps in the delivered format (for example RLDS, HDF5 or MCAP).
- Task coverage: object list, scene count, lighting conditions, clutter levels, and the share of episodes containing failures and recoveries.
- Episode metadata: operator ID (pseudonymized), success label and who assigned it, language instruction if any, and calibration files per session.
- Evaluation split: a held-out subset by scene and object, never by random episode, so the split tests generalization.
- Rights: ownership chain, worker and bystander consent, permitted uses, and asset licenses for any synthetic augmentation the supplier adds.
Before paying, run the acceptance checks for robot datasets, and review what drives robot data cost to see which line items move price.
Where SourceX fits in a real-data plan
SourceX helps buyers source the real-world side of the mix, not simulation. It sources operational datasets from US companies on request, including new recordings of hands-on work, and manages the licensing process; it does not hold stock, so a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, with personal details removed or replaced before delivery. SourceX does not train models. For the wider category, see the multimodal and embodied data hub, robotics training data for embodied AI, our guide to combining licensed and synthetic data, and the definitions of synthetic data and real-world data. Teams that already know which sub-skills need physical data can describe the request to SourceX.
Sourcing real robot demonstrations for a co-training mix
If your worksheet shows sub-skills where simulation falls short, describe the data you need: embodiment, sensors, tasks and evaluation split. SourceX looks for US businesses that hold the described data, assesses data and licensing permissions, and nothing is contracted until a supplier agrees. Describe the real robot data you need.
Sources
- Open X-Embodiment Collaboration (arXiv), "Open X-Embodiment: Robotic Learning Datasets and RT-X Models" (2023). https://arxiv.org/abs/2310.08864v1
- Khazatsky et al. (arXiv), "DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset" (2024). https://arxiv.org/abs/2403.12945v2
- arXiv, "Data Pyramid for Embodied Manipulation: A Survey" (2026). https://arxiv.org/pdf/2607.24744
- arXiv, "Phantom: Training Robots Without Robots Using Only Human Videos" (2025). https://arxiv.org/pdf/2503.00779
- Ou et al., NeurIPS 2024 (arXiv), "Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale" (2024). https://arxiv.org/pdf/2409.15637
- Mayer Brown, "Synthetic data as a deal asset: ownership, provenance and diligence considerations in AI acquisitions" (2026). https://www.mayerbrown.com/en/insights/publications/2026/07/synthetic-data-as-a-deal-asset-ownership-provenance-and-diligence-considerations-in-ai-acquisitions
- Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Pushkarna, Zaldivar, Kjartansson, Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.