Skip to content

Multimodal and embodied data

Acceptance Checks for Robot Datasets Before You Pay

Quick answer

Robot dataset quality checks should be objective tests you run on a sample, and again on the full delivery, before payment: every declared stream present per episode, timestamps synchronized within a stated tolerance, calibration files that reproject correctly, actions that line up with observed motion, success labels that survive independent re-labeling, and diversity that matches the specification. Agree the thresholds, the sampling plan and the remedy in writing before the first file arrives.

By SourceX Editorial · Updated

Why robot data needs its own acceptance tests

Robot demonstrations fail in ways generic quality reviews miss, because the defects live between streams rather than inside any one of them. A wrist-camera video can look perfect while lagging joint states by 80 ms, and a gripper channel can be complete but scaled in the wrong units. Some buyers want an independent verdict before releasing funds, and that verdict is offered as a stand-alone audit service [1].

The tests below assume you already specified the data up front, as covered in what to specify when sourcing teleoperation demonstrations. This page covers the other half: proving a delivery meets that specification. For the contractual mechanics of inspection windows, rejection notices and cure periods, see dataset acceptance testing process.

Stage the checks: evaluation sample first, full delivery second

Run the full battery twice, first on an evaluation sample and then on the delivery, so you discover format and sync problems before the supplier has produced thousands of hours. Sample access can be granted under an evaluation-only license that limits use to evaluation rather than production, as in NVIDIA's sample data license [2]; confirm your sample license allows the specific tests you plan, including replaying actions on hardware.

A practical split:

  • Sample stage (tens of episodes): schema, stream completeness, sync, calibration and replay. These are cheap and catch systematic defects.
  • Delivery stage (all episodes, automated; a random subset, manual): completeness and sync on everything, success-label re-labeling and diversity on a statistically drawn subset.
  • Lot structure: if the supplier ships in batches, treat each batch as a lot and use an attribute sampling plan such as ANSI/ASQ Z1.4, which defines AQL-based plans with tightened, normal and reduced inspection [9].

The guide to running a data pilot with a supplier covers how to scope the sample stage commercially.

Schema and stream completeness

Completeness is the first gate: every episode must contain every declared stream for its full duration, or the episode is incomplete. Parse the container before you look at content. In RLDS, data is organized as episodes made of steps, each step carrying observation, action and reward fields plus boundary flags such as is_first, is_last and is_terminal [3]; in MCAP or ROS 2 bag files, data arrives as channels of timestamped messages [4].

Checks to automate:

  • Every episode has the topics or feature keys listed in the spec (for example, three camera streams, joint positions, joint velocities, gripper command, end-effector pose, language instruction).
  • Exactly one is_first and one is_last per RLDS episode; no truncated episodes flagged as terminal.
  • Per-stream message counts match nominal rate times duration within a tolerance (a 30 Hz camera over 20 s should yield roughly 600 frames).
  • No silent dtype or unit drift: joint angles in radians throughout, gripper width in meters or normalized consistently.
  • Gaps: compute the longest inter-message interval per stream and flag anything above, say, three nominal periods.

Open X-Embodiment pooled datasets from many labs into one RLDS-based format, and the conversion work it describes is a reminder that heterogeneous action spaces and coordinate conventions are a normal source of error [5].

Timestamp synchronization across cameras and proprioception

Synchronization passes when the offset between each stream and the reference clock stays inside the tolerance you set, measured on content, not only on header timestamps. MCAP records both a log time and a publish time per message [4]; a large or drifting gap between them suggests buffering or clock problems on the recording machine.

Header timestamps can be wrong while looking consistent, so add a content test. Find events visible in several streams (gripper closing on camera and in the gripper width channel, a contact spike in the force-torque signal) and measure the lag between them across a sample of episodes. Report the median and the 95th percentile offset; a stable offset can be corrected, but a drifting one usually cannot. For contact-rich data, where sync tolerances are tighter, see tactile and force-torque data.

Calibration validity

Calibration passes when intrinsics and extrinsics supplied with each episode reproduce known geometry in the images. Datasets such as DROID treat calibrated camera poses as a first-class dataset property [6], and a buyer should expect calibration files per camera and per session rather than one global file.

Tests:

  • Reprojection: project the end-effector position from forward kinematics into each calibrated camera and measure pixel error against the visible gripper. Large or session-dependent errors mean extrinsics drifted or were copied between sessions.
  • Session coverage: every episode references a calibration record with a timestamp from the same session.
  • Depth sanity: if depth is delivered, check depth at the projected gripper location against the kinematic distance.

Action-observation alignment and replay

Action alignment passes when commanded actions explain observed motion, and the strongest test is replaying recorded actions on a reference robot or a faithful simulator. Integrate commanded joint velocities or deltas and compare against measured joint positions; persistent residuals reveal off-by-one-step errors, wrong action frames (base versus end-effector) or a controller mismatch between what was logged and what was executed.

Replay a sample of episodes open-loop on the same robot model and controller. Exact task success will not reproduce because object positions differ, but the arm trajectory should track within your stated tolerance. If the supplier recorded with a different gripper or firmware than the spec states, this is where it shows.

Success-label verification

Success labels pass when an independent re-labeling of a random subset agrees with the supplier above a pre-agreed rate. Label noise is common even in well-known datasets; one audit estimated an average test-set label error rate of at least 3.3% across ten benchmarks [8]. Success detection itself is a hard classification problem whose accuracy has to be measured against human judgment [7].

Procedure:

  1. Write a success definition per task that a reviewer can apply from video alone (for example, "mug upright on the coaster, gripper open, arm retracted").
  2. Draw a stratified random sample across tasks, operators and sessions.
  3. Have two reviewers label blind to the supplier label; adjudicate disagreements.
  4. Report agreement with the supplier label, Cohen's kappa between your reviewers, and the error direction (false successes are usually more damaging than false failures for imitation learning).

Episodes labeled as failures or interventions need the same treatment if you plan to use them; see robot failure, intervention and recovery data. General annotation audit methods are in how to audit annotation quality.

Diversity and coverage against the specification

Diversity passes when measured counts of scenes, objects, operators, lighting conditions and initial states meet the minimums in your spec, computed from metadata and spot-checked against video. DROID, for example, reports its scene and task breadth as headline properties [6]; ask for the same tables for a commercial delivery and verify them.

Compute per-axis histograms and flag concentration: if one operator produced most episodes, or one scene dominates a task, the dataset may teach that operator's habits. Near-duplicate detection on first frames catches episodes recorded from an identical reset.

Acceptance criteria worksheet

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckMeasured onMetricExample thresholdIf it fails
Stream completenessAll episodesEpisodes with every declared stream, full duration≥ 99%Reject affected episodes; supplier replaces
Message rateAll episodesCount vs nominal rate × durationWithin ±2%Flag stream; investigate recorder
Sync offset50 sampled episodesp95 cross-stream event lag≤ 33 ms (one 30 Hz frame)Request corrected timestamps
Calibration50 sampled episodesMedian gripper reprojection error≤ 10 pxRe-calibrate or drop sessions
Action replay20 episodes on reference robotMean joint tracking errorWithin agreed toleranceHold payment pending explanation
Success labelsStratified 200 episodesAgreement with supplier label≥ 95%; kappa ≥ 0.8 between reviewersRe-label at supplier cost
DiversityMetadata + video spot checkScenes, objects, operators vs specMeets each minimumAdditional collection or price adjustment

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "episode_id": "ep_000412",
  "session_id": "s_2026_09_14_a",
  "robot": {"model": "7-dof arm", "gripper": "parallel jaw", "controller": "joint_velocity_20hz"},
  "streams": {"cam_wrist": 30, "cam_left": 30, "joint_pos": 100, "gripper": 100, "ft_sensor": 500},
  "calibration_ref": "calib/s_2026_09_14_a.yaml",
  "task": "place_mug_on_coaster",
  "instruction": "put the mug on the coaster",
  "operator_id": "op_07",
  "success": true,
  "success_labeler": "supplier_qc",
  "duration_s": 21.4
}

A manifest like this lets you compute completeness, calibration coverage, label provenance and diversity without opening a single video.

Rights checks belong in the same acceptance gate

Technical acceptance is not enough if the episodes cannot be used as licensed. Confirm that operator consents cover the recordings (faces and voices often appear in third-person views), that the facility owner and robot operator have agreed to release, and that any open-source components carry compatible licenses; see who owns robot data and checking open robot dataset licenses. Documenting the tests and results also fits the MEASURE function of the NIST AI RMF if your organization uses it [10].

The broader evaluating data supplier quality guide covers supplier-level questions, and the multimodal and embodied data hub links related robot data topics.

Where SourceX fits

SourceX sources operational datasets from US companies, including new recordings of hands-on work, on request rather than from stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Buyers can describe the robot data they need and the acceptance tests it must pass; the robotics and embodied AI use case explains the kinds of data involved.

Get robot demonstration data that passes your acceptance checks

SourceX looks for US businesses that hold the data you describe and manages the process from assessment of data and licensing permissions through a license that defines records, uses, term and delivery. Nothing is contracted until a supplier agrees. Start by telling SourceX what robot data you need.

Frequently asked questions

How large should the success-label re-labeling sample be?

Size it for the precision you need on the agreement rate, stratified across tasks and operators. A few hundred episodes gives a usable estimate for a single pass/fail threshold; for lot-by-lot deliveries, an attribute sampling plan such as ANSI/ASQ Z1.4 sets sample sizes by lot size and AQL [9].

Can we accept a dataset without a physical reference robot?

Yes, but replace hardware replay with kinematic consistency checks and replay in a simulator using the same robot description and controller. State in the acceptance criteria that simulator replay is the agreed test, so neither side disputes it later.

What should happen to episodes that fail?

Decide in advance whether failing episodes are dropped, replaced or repaired, and whether payment is per accepted episode or per lot. The dataset acceptance testing process page covers rejection notices and cure periods.

Sources

  1. Contra (service listing), "Robot dataset acceptance audit: a verdict before you pay". https://contra.com/s/ee6vvEQU-robot-dataset-acceptance-audit-a-verdict-before-you-pay
  2. NVIDIA, "NVIDIA Sample Data License for Evaluation (2026.01.19)" (2026). https://developer.download.nvidia.com/licenses/nvidia-sample-data-license-for-evaluation-2026.01.19.pdf
  3. Ramos et al. (Google Research), "RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning" (2021). https://ar5iv.arxiv.org/html/2111.02767
  4. Kaitai Struct format gallery, "MCAP: format specification". https://formats.kaitai.io/mcap/index.html
  5. Open X-Embodiment Collaboration, "Open X-Embodiment: Robotic Learning Datasets and RT-X Models" (2023). https://arxiv.org/pdf/2310.08864
  6. Khazatsky et al., "DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset" (2024). https://arxiv.org/pdf/2403.12945
  7. Du et al. (DeepMind), "Vision-Language Models as Success Detectors" (2023). https://arxiv.org/pdf/2303.07280
  8. Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  9. ASQ, "ASQ/ANSI Z1.4:2003 (R2018): Sampling Procedures and Tables for Inspection by Attributes". https://asq.org/quality-press/display-item?item=T1164
  10. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data