Skip to content

Multimodal and embodied data

Robot Failure, Intervention and Recovery Data

Quick answer

A useful robot failure dataset records what went wrong, when it went wrong, whether it was recoverable, and what a human or the system did next. Success-only demonstrations teach the nominal path; failure and intervention data teach the states a policy actually drifts into. When sourcing, ask for episodes with a labeled failure onset, an intervention record (who took over, through which device, for how long) and the post-takeover corrective trajectory, plus the share of failed episodes in the set.

By SourceX Editorial · Updated

Why success-only demonstrations leave policies brittle

Success-only demonstrations leave policies brittle because a behavior-cloned policy never sees the off-nominal states its own small errors create. Large open corpora such as DROID are built from teleoperated demonstrations across many scenes [3], and pooled collections such as Open X-Embodiment standardize many such datasets into one episode format [1]; both are valuable, but they are organized around task attempts, not around labeled failures and takeovers. Interactive imitation learning addresses the gap by collecting expert corrections in the states the learner actually visits. In human-gated variants, the operator takes over control when the policy heads toward a failing state, so the corrective actions come from someone driving the robot rather than from after-the-fact relabeling.

Failure records also matter for reinforcement learning. Work on recovery from tool-execution errors argues that standard RL reduces a failure to a sparse negative reward that gives no guidance on how to recover [7]; the same logic applies to manipulation, where a recovery trajectory carries far more signal than a "task failed" flag.

The same bias affects evaluation. A held-out set drawn from curated successful teleoperation runs measures how well a policy imitates good runs, not how it behaves after a slipped grasp or a misdetected object. If your evaluation set has no failures, your failure-detection classifier has no positives and your success-rate estimate has no denominator for recovery. For broader framing on how demonstration data is specified, see robot teleoperation data.

Three record types: failures, interventions and recoveries

Failure, intervention and recovery data are three distinct record types, and a supplier that has one often lacks the others. Keep them separate in your request so you can price and accept each on its own terms.

  • Failure episodes. Autonomous or teleoperated runs that ended without task success or crossed a safety limit. Value comes from the failure onset label and the sensor context around it (RGB-D, wrist camera, joint states, gripper width, force-torque, audio).
  • Intervention segments. Moments where a human operator took control from the policy, via a teleoperation device, joystick, spacemouse or e-stop. Value comes from the trigger, the handover timestamp and the operator's actions during control.
  • Recovery demonstrations. Trajectories that start from an off-nominal state (a dropped part, a misaligned peg, a jammed drawer) and return the task to a recoverable state or complete it. These may be collected deliberately by staging failures, or extracted from intervention segments.

Some research failure benchmarks mix simulated failures injected into task executions with a smaller set of real, teleoperated failures, and split planning failures from execution failures. That split matters commercially: planning failures (wrong object, wrong order) are often labeled from video and task logs, while execution failures (slip, collision, missed grasp) need proprioception and contact signals to label well. If contact matters, read our guide to tactile and force-torque data.

A failure and intervention label taxonomy

A workable taxonomy labels failure type, onset time, recoverability and intervention details at the episode and segment level. The fields below are a starting point; adjust classes to your task family rather than adopting a generic list.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "episode_id": "ep_000418",
  "task": "bin_pick_to_tote",
  "robot": {"arm": "6-dof", "gripper": "parallel_jaw", "control_hz": 15},
  "policy_version": "pi_2026_08_r3",
  "outcome": "failed_then_recovered",
  "failure": {
    "class": "execution",
    "subtype": "grasp_slip",
    "onset_t_s": 12.40,
    "detected_t_s": 13.05,
    "detected_by": "operator",
    "recoverable": true,
    "severity": "no_damage",
    "annotator_confidence": 0.8
  },
  "intervention": {
    "trigger": "operator_judgment",
    "device": "spacemouse",
    "start_t_s": 13.10,
    "end_t_s": 19.75,
    "operator_role": "trained_teleoperator",
    "control_mode": "full_takeover"
  },
  "recovery": {
    "strategy": "regrasp_after_reorient",
    "returned_to_policy": true,
    "task_success": true
  },
  "streams": ["rgb_wrist", "rgb_scene", "depth_scene", "joint_pos", "joint_vel", "gripper_width", "ft_wrist"],
  "format": "rlds_episode"
}

Define each field operationally before labeling. The difference between onset_t_s and detected_t_s is the detection latency your failure detector must beat, and it is lost if annotators mark only the moment of takeover. Recoverability should be judged from the state, not the outcome: an episode can be recoverable and still fail because no one intervened. Record trigger as one of operator judgment, automated safety limit (joint torque, workspace boundary, force threshold), uncertainty gate, or scheduled check, since some interactive learning systems gate takeovers on model uncertainty and you may want to compare gate types. If you consume RLDS-style episodes, carry these labels as episode and step metadata rather than in side files, so they survive conversion [2].

Common labeling failure modes to check for:

  • Takeover equals failure. Annotators label every intervention as a failure, although some operators intervene early out of caution. Ask for a "precautionary" flag.
  • Post-hoc onset. Onset is marked at the e-stop rather than at the first divergent state, which shortens every failure window.
  • Clipped context. Logs start at the intervention, so the pre-failure states the policy must learn to recognize are missing. Request a fixed pre-roll (for example several seconds) per segment.
  • Collapsed classes. "Other" absorbs most failures. Track the share of "other" per batch and require subclassing above an agreed level.
  • Unsynchronized streams. Teleoperation commands logged on a different clock from cameras. Ask for the timestamp source and measured offset, the same acceptance logic described in acceptance checks for robot datasets.

Which statistics to request before licensing

Request summary statistics first, because the value of a failure dataset depends on mix, not volume. A request built around these numbers lets a supplier answer from logs without sharing raw episodes.

StatisticWhy it mattersTypical red flag
Failure share of episodesDetermines how many positives a detector sees and how biased a success rate will beShare near zero, or failures removed during "cleaning"
Failure class distributionShows whether the set covers your dominant failure modesOne class (often grasp failure) dominates
Intervention rate per hour of autonomyIndicates how much corrective signal existsInterventions logged only as e-stop events with no operator actions
Median intervention durationShort takeovers rarely contain a full recoveryMost takeovers end in reset, not recovery
Recovery success after takeoverSeparates useful corrections from abandoned episodesRecoveries not linked to the failure that caused them
Policy versions coveredInterventions are policy-specific; old versions fail differentlyNo policy version field
Operator count and trainingCorrection style varies by operatorSingle operator, no role field

Interventions are tied to the policy that was running, so corrections collected against an old policy target states your current policy may never reach. Ask for the policy version on every segment and weight recent versions accordingly. If the source is a production deployment, see operational logs from deployed robot fleets for how those logs are usually structured.

Where failure data comes from and what each source misses

Each source of failure data has a characteristic gap, so most buyers combine two. The table compares the usual routes.

SourceStrengthGap to check
Deployed fleet logsReal failure distribution and real operatorsRetention may keep only flagged clips; safety incidents may be withheld
Interactive imitation learning collection (human-gated)Corrective actions aligned to the current policyNarrow to one policy and task set
Staged failure and recovery sessionsControlled coverage of rare classesFailures may look scripted; check onset realism
Simulation with injected failuresCheap, labeled ground truthSim-to-real gap on contact and perception
Human-only video of mistakesScale for planning-level errorsNo actions or proprioception
Open demonstration corpora [1][3]Broad scenes and embodimentsIncludes unsuccessful episodes; check license terms before commercial use

For the cost and sim tradeoffs, see real-world vs simulated robot data and what drives the cost of robot training data. For language-paired failure explanations, the vision-language-action training data guide covers instruction pairing.

Screening safety incidents, operator identity and rights

Failure data needs more screening than success data because failures overlap with safety incidents, operator performance and customer sites. Screen these before any license is drafted.

First, safety incidents. Episodes involving contact with a person, property damage or a near miss may sit in incident reports, insurer files or internal investigations that a supplier is not free to release. Ask the supplier to confirm which incident categories are excluded and whether exclusion changes the failure distribution you receive. Second, operator identity. Intervention data often includes operator IDs, voice from headsets, and hands or faces in scene cameras; Texas, for example, treats records of hand or face geometry as biometric identifiers requiring notice and consent before capture for a commercial purpose [5]. Third, purpose. Logs collected for maintenance or quality control under customer terms may not cover AI training, and the FTC has warned that retroactively broadening data use for AI training may be unfair or deceptive [6].

Ownership is often split between the robot operator, the OEM and the facility, which is covered in who owns robot data. De-identifying faces, voices and screens together is covered in de-identifying multimodal records. For risk documentation on how you used failure data in evaluation, the NIST AI RMF MEASURE and MANAGE functions give a structure many enterprise reviewers already recognize [4].

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

How SourceX handles requests for robot failure data

SourceX sources operational datasets from US companies on request and manages the commercial process, including the license; it does not hold inventory, and a request does not guarantee a match. Buyers describe the data they need, such as the taxonomy and statistics above, and SourceX looks for US businesses that hold it; every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows only after an executed agreement. You can describe your failure and intervention data needs as a starting point, and see robotics training data for embodied AI and training data for RL environments for related requests. The multimodal and embodied data hub links the rest of this cluster, and the AI data guides cover other modalities.

Request failure, intervention and recovery data

SourceX sources datasets on request from US companies, with rights review and a license defining records, uses, term and delivery for each deal. Nothing is contracted until a supplier agrees, and terms are agreed per deal. Share your taxonomy, target failure share and streams at sourcex.si/buyers.

Frequently asked questions

Is intervention data the same as a failure dataset?

No. An intervention marks when a human took control, which may be precautionary, while a failure label marks that the task or a safety limit actually failed. Label both and keep a precautionary flag so a detector does not learn operator caution as failure.

How much pre-failure context should each segment include?

Enough to cover the states a detector must recognize before onset; agree a fixed pre-roll per segment in the request. Segments that begin at the takeover remove exactly the signal a failure predictor needs.

Can simulated failures replace real ones?

They help with labeled coverage of rare classes, but contact dynamics and perception errors differ in the real world. Use simulation to fill classes, and real episodes to calibrate rates and evaluate.

Sources

  1. Open X-Embodiment Collaboration (arXiv:2310.08864), "Open X-Embodiment: Robotic Learning Datasets and RT-X Models" (2023). https://arxiv.org/html/2310.08864v8
  2. Google Research (arXiv:2111.02767), "RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning" (2021). https://ar5iv.arxiv.org/html/2111.02767
  3. DROID authors (arXiv:2403.12945), "DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset" (2024). https://arxiv.org/html/2403.12945v2
  4. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
  5. Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier" (2026). https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  6. Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  7. Zhang et al. (arXiv:2601.15625), "Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors" (2026). https://arxiv.org/html/2601.15625v1

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data