Skip to content

Schemas, packaging and delivery

Robotics and Sensor-Fusion Data Formats: MCAP, ROS Bags, RLDS and LeRobot

Quick answer

For robot and teleoperation data, ask for two layers. The raw layer should be MCAP (or the original rosbag2 files), because it keeps every topic, message schema and nanosecond timestamp exactly as recorded [1][2]. The training layer should be RLDS or LeRobotDataset, which resample those streams into episodes and steps that loaders can batch [3]. Specify clock sources, calibration files and the conversion script, because most information loss happens when the raw layer is converted into the training layer.

By SourceX Editorial · Updated

Raw logs versus training-ready episodes

A raw log is a faithful record of a robot's message bus, while a training format is a lossy, opinionated view of it. Raw logs hold asynchronous streams: a wrist camera at 30 Hz, joint states at 500 Hz, a force-torque sensor at 1 kHz, /tf transforms and operator commands, each on its own timeline. Training formats force those streams onto one step clock and keep only the fields someone chose. If you receive only the training layer, you cannot re-cut episodes, change the control rate, or add a sensor stream later.

This page covers formats. For what to specify about the demonstrations themselves (operators, task design, success labels), see robot teleoperation data. The delivery cluster hub covers manifests, transfer and versioning across all data types.

MCAP and ROS bags: the recording layer

MCAP is a self-describing container for timestamped pub/sub messages, and recent ROS 2 releases can store bags in it through the rosbag2_storage_mcap plugin [2]. An MCAP file holds Schema records (for example a ros2msg definition of sensor_msgs/msg/JointState), Channel records that bind a topic to a schema and message encoding (cdr, protobuf, json), and Message records carrying the payload [1]. Optional Chunk, MessageIndex, Statistics and Summary sections allow seeking by time without reading the whole file.

Each Message carries log_time and publish_time as unsigned nanosecond counts from an epoch the recorder chooses, such as Unix time or robot boot time [1]. That flexibility is the most common failure: two bags recorded on different machines can share field names but use different epochs. Ask the supplier which clock each topic used and whether hosts were synchronized with PTP (IEEE 1588) or NTP.

Older material arrives as ROS 1 .bag files or rosbag2 SQLite3 .db3 files with a metadata.yaml. Both are readable, but .db3 bags from older ROS 2 distributions often lack the embedded message definitions that MCAP carries, so you also need the exact message package versions. Custom message types are where undocumented deliveries break.

RLDS: episodes and steps for TensorFlow pipelines

RLDS stores data as a dataset of episodes, each containing a sequence of steps, and it was designed to preserve sequential decision-making information that flat transition tuples lose [3][4]. A step typically carries observation (a nested dict of images, proprioception and language), action, reward, discount, and the boundary flags is_first, is_last and is_terminal [3]. Episode-level metadata (robot ID, task, success, operator) sits alongside the steps.

RLDS datasets are usually built as TensorFlow Datasets builders and serialized as TFRecord shards [4]. The Open X-Embodiment collaboration converted datasets from many robots and labs into a common RLDS-based format, which is why many vision-language-action codebases still expect it [5]. Image observations are usually stored as encoded frames per step, so files grow quickly with camera count and resolution.

The key difference from a raw log: is_terminal distinguishes a true task end from a time-limit cut, which matters for value learning. Ask suppliers how they set it, because some converters simply copy is_last into it.

LeRobotDataset: Parquet tables plus video

According to the LeRobot project documentation (verify against the lerobot release you use), LeRobotDataset v3 stores per-frame numeric data in chunked Parquet files and camera streams as MP4 video, with schema and episode boundaries in metadata files. The layout has data/ (Parquet with many episodes per file), videos/ (MP4 shards per camera key), and meta/ holding info.json, normalization statistics, task descriptions and per-episode records under meta/episodes/. info.json declares features with shapes and dtypes, fps, codebase_version and the path templates for data and video shards.

Parquet's footer metadata lets a loader read only the columns it needs, for example observation.state and action, without touching other columns [6]. Video compression keeps multi-camera datasets far smaller than per-frame images, at the cost of codec artifacts and decode-time frame seeking. Check codebase_version rather than a dataset's title, since v2.x and v3.0 layouts differ and v3 requires a recent lerobot release.

Comparison: which format for which job

The right choice depends on whether you need fidelity, a specific training stack, or both.

Illustrative example: invented to show structure; it does not describe an available dataset.

QuestionMCAP / rosbag2RLDS (TFDS, TFRecord)LeRobotDataset v3
Primary roleLossless recording and replayEpisode/step training dataEpisode/step training data with video
Time modelPer-message nanosecond log_time, publish_time [1]One step index per episodeFixed fps, frame index and timestamp per row
SchemasEmbedded per channel [1]TFDS feature specfeatures in meta/info.json
ImagesRaw or compressed message payloadsEncoded frames per stepMP4 per camera
Tooling fitROS 2, Foxglove, custom replayTensorFlow, JAX, Open X-Embodiment code [5]PyTorch, Hugging Face Hub
Main riskUnknown epochs, missing custom msgsLost raw rates, ambiguous is_terminalCodec artifacts, version drift

Synchronizing camera and joint timestamps

Synchronization is a delivery requirement, not a post-processing detail, so specify the method in the request. A camera frame has at least three times: exposure (sensor capture), driver receipt, and log write. Joint encoders are sampled on the controller's clock. When a converter aligns by log-write time, a 20 to 60 ms camera pipeline delay turns into a systematic offset between what the robot saw and what it did.

Ask for the following, recorded per episode:

  • Clock source per topic (hardware trigger, PTP, NTP, host monotonic) and the epoch used in log_time [1].
  • The alignment rule used when resampling: nearest-neighbor, previous-sample hold, or interpolation, with the tolerance window in milliseconds.
  • Dropped-frame and gap counts per stream, not just overall episode length.
  • Whether actions are commanded targets or measured states, and their lag relative to observations.

See timestamps and time zones in delivered datasets for ordering rules that apply to all event data.

Calibration and metadata that conversions drop

Training formats often discard the calibration and frame metadata that make sensor-fusion data reusable. Intrinsics (sensor_msgs/CameraInfo: K, D, distortion model), camera-to-base extrinsics from static /tf_static, URDF version, gripper type and joint limits rarely survive RLDS or LeRobot conversion unless someone adds them as episode metadata. For camera, LiDAR and radar rigs, see sensor-fusion data specification.

Common conversion losses to check:

  • Rate loss: 1 kHz force data averaged or dropped to a 10 to 30 Hz step rate.
  • Frame loss: /tf trees flattened into a single end-effector pose, with the reference frame unnamed.
  • Unit drift: radians versus degrees, quaternion order (xyzw versus wxyz), gripper width versus normalized 0 to 1.
  • Episode boundaries: operator resets and failed attempts silently trimmed.
  • Image changes: resizing, cropping or re-encoding without recording parameters.

Delivery specification template

A written specification turns format choices into checkable acceptance criteria. Adapt the template below for a request or a license schedule.

Illustrative example: invented to show structure; it does not describe an available dataset.

delivery_spec:
  raw_layer:
    format: mcap            # or rosbag2 .db3 with metadata.yaml
    message_encoding: cdr
    include_custom_msg_packages: true
    clock_per_topic_documented: true
  training_layer:
    format: lerobot_v3      # or rlds_tfds
    fps: 30
    features: [observation.images.wrist, observation.images.front,
               observation.state, action]
    action_semantics: commanded_joint_positions
    alignment_rule: previous_sample_hold
    alignment_tolerance_ms: 15
  per_episode_metadata: [robot_model, urdf_version, gripper, task_text,
                         success, operator_id_pseudonymous, calibration_ref]
  calibration: [camera_info_yaml, extrinsics_tf_static, urdf]
  conversion: {script_repo_commit: required, raw_to_train_episode_map: required}
  verification: {checksums: sha256_manifest, sample_replay: 10_episodes}

Pair it with a checksum manifest and a machine-readable dataset description; Croissant's JSON-LD vocabulary can describe file sets and record fields for both layers [7]. Our guide to Croissant metadata for licensed datasets lists the fields worth requesting, and video dataset packaging covers codec and sidecar choices.

Licensing and privacy checks specific to robot logs

Robot logs are operational records, so format work runs alongside rights and privacy review. Cameras in workplaces capture people, screens and documents, and teleoperation logs can contain operator IDs, so faces, names and account details in frames or metadata need handling before delivery. Fleet logs from production sites raise similar questions; see operational logs from deployed robot fleets.

SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and manages the licensing process with the supplying company. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Buyers can describe the robot or sensor data they need by format and fields, and the robotics and embodied AI use-case page and sensor and IoT data page cover the data needs.

Sourcing robotics data in the format you specify

SourceX looks for US businesses that hold the data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered through private, access-controlled workflows. Data is sourced on request, so a request does not guarantee a match, and nothing is contracted until a supplier agrees. Share your format and field specification with SourceX.

Sources

  1. MCAP Format Specification, "MCAP: format specification". https://formats.kaitai.io/mcap
  2. Open Robotics Discourse, "MCAP: A recording file format for pub/sub systems, designed with robotics in mind". https://discourse.openrobotics.org/t/mcap-a-recording-file-format-for-pub-sub-systems-designed-with-robotics-in-mind/24315
  3. Ramos et al., arXiv:2111.02767, "RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning" (2021). https://ar5iv.arxiv.org/html/2111.02767
  4. Google Research, "RLDS: An Ecosystem to Generate, Share, and Use Datasets in Reinforcement Learning" (2021). https://research.google/blog/rlds-an-ecosystem-to-generate-share-and-use-datasets-in-reinforcement-learning/
  5. Open X-Embodiment Collaboration, arXiv:2310.08864, "Open X-Embodiment: Robotic Learning Datasets and RT-X Models" (2023). https://arxiv.org/abs/2310.08864v1
  6. The Apache Software Foundation (Apache Parquet), "File Format". https://parquet.apache.org/docs/file-format/
  7. Akhtar et al. (MLCommons), arXiv:2403.19546, "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data