Skip to content

Multimodal and embodied data

Open Robot Learning Datasets: Checking Licenses Before Commercial Training

Quick answer

Open robot datasets can sometimes be used for commercial training, but a pooled collection's top-level license rarely settles the question. Mixtures such as Open X-Embodiment combine subsets from many labs, so you have to trace the terms of each subset you load, confirm who owns the recordings, and check consent for people and facilities captured on camera. Treat the hub license tag as a lead, not an answer, and keep a per-subset audit record before any weights ship.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why pooled robot datasets need a per-subset license check

A pooled robot dataset is a wrapper around many contributors' releases, and each contributor may have set its own terms. Open X-Embodiment, for example, assembles over one million trajectories from 22 robot embodiments contributed by many institutions [1]. Any umbrella license line on such a mixture, often Apache 2.0 for code and a Creative Commons license for other materials, sits at the repository level, while the data itself is a collection of separately produced constituent datasets.

The practical consequence is that the umbrella statement is a starting hypothesis. A constituent dataset may have been published first on a lab website, in a paper appendix or on a project page with different wording, and the pooled release may not restate it. Your counsel needs the upstream text for each subset you actually train on, not the license of the conversion repo.

Single-lab releases are simpler but not trivial. DROID, for instance, was released by its authors as an open-source dataset with three synchronized RGB camera streams, calibration, depth and language instructions per episode, gathered across many scenes and buildings [2]. Even there, confirm the dataset license on the project's own release page rather than relying on the license shown for the paper on arXiv, which covers the article text.

Why hub license fields are not enough

Hub license fields are frequently missing or wrong, so you should trace each component back to its original terms. The Data Provenance Initiative audited more than 1,800 text datasets and reported license omission above 70% and license error rates above 50% on popular hosting sites [3]. Robot data was not its subject, but the mechanism is the same: whoever uploads a mirror chooses the tag.

On Hugging Face, the dataset card and its YAML metadata, including the license tag, are written by the uploader through the Hub metadata workflow [4]. A community conversion of an RLDS subset into LeRobot or another format may carry license: apache-2.0 or cc-by-4.0 copied from a parent repo. That tag tells you what the uploader believed, not what the original lab granted.

Common failure modes when auditing open robot data:

  • Repo license applied to data. Apache 2.0 on loader code read as a data grant.
  • Paper license applied to data. CC BY on the arXiv article read as a dataset license.
  • Mirror drift. A re-hosted copy drops a non-commercial clause or an attribution requirement.
  • Version mismatch. Terms changed between releases, and your shard came from an older or newer version than the one reviewed.
  • Bundled assets. Object meshes, CAD models, scene scans or pretrained encoders inside the release carry third-party terms.

For the general method across text, image and audio corpora, see the open dataset license audit and the open data license compatibility matrix.

Robot recordings raise consent and ownership questions that a copyright license does not answer. Third-person and wrist cameras capture operators' hands, faces and bodies, bystanders, home interiors, labs and workplaces. A license from a lab covers what the lab could grant; it does not retroactively supply consent from the people filmed.

Biometric statutes are the sharpest edge. Texas Business and Commerce Code Section 503.001 defines biometric identifiers to include records of hand or face geometry and bars capturing them for a commercial purpose without informing the individual and obtaining consent first [5]. Whether a given video stream amounts to such a record is a question for counsel, but face-visible teleoperation footage is the kind of material that prompts it.

Facility and equipment owners can also hold interests. Recordings made in a company warehouse or with a vendor's robot may be subject to site agreements or OEM terms. The robot data ownership guide covers how operators, OEMs and facility owners split those rights, and multimodal de-identification covers blurring faces and screens across synchronized streams.

Regulatory documentation that depends on your audit

Your audit trail feeds disclosure duties, so build it in the shape regulators and customers will ask for. Article 53 of the EU AI Act requires providers of general-purpose AI models to keep a policy for complying with Union copyright law and to publish a summary of training content [7]. These duties have applied since 2 August 2025, with AI Office enforcement powers from 2 August 2026, as of October 2026.

The Commission published its template for the public training-content summary on 24 July 2025 [8]. Whether a robot foundation model counts as a general-purpose AI model is a scoping question for counsel, but teams that keep per-subset provenance can answer it either way without reconstructing history.

In the US, the Copyright Office's Part 3 report on generative AI training remains a pre-publication version as of October 2026 and analyzes how training can implicate copyright [6]. It is not binding law, and courts are still deciding the issues, so the defensible position is a clean record of what you used and under which terms.

Per-subset license audit record

The core artifact is one record per constituent subset, completed before that subset enters a training mixture. Keep it in version control next to the data-loading config so the mixture definition and the rights record cannot drift apart.

Illustrative example: invented to show structure; it does not describe an available dataset.

subset_id: oxe/example_lab_kitchen_v1      # name as used in the mixture config
loaded_from: rlds mirror, version 1.0.0     # exact source and version you pulled
upstream_release: lab project page, 2023-xx # original publisher of the data
license_text_location: upstream README + paper data statement
license_as_written: "CC BY 4.0"             # quoted, not inferred from a hub tag
hub_tag_matches_upstream: false             # mirror showed apache-2.0 (code license)
commercial_training_permitted: yes, with attribution
attribution_requirement: cite paper + lab in model card
share_alike_or_nc_terms: none found
bundled_third_party_assets: object meshes (separate terms, not used)
people_visible: operator hands, occasional faces
consent_evidence: none published; flagged for counsel
facility_type: university lab
hardware: Franka Panda, wrist + 2 external RGB
decision: include, faces blurred before training
reviewer: counsel initials, date

Fields that most often change the decision are license_as_written, share_alike_or_nc_terms, people_visible and consent_evidence. A subset with an NC clause, an unclear license or no consent story should drop out of the commercial mixture, even if it remains usable for internal non-commercial research where your counsel agrees.

Decision table: what to do with each subset

Every subset should land in one of a few outcomes, and the outcome should be recorded rather than left implicit.

FindingTypical actionWatch for
Permissive license (e.g., CC BY 4.0) traced to upstream textInclude; meet attribution termsBundled assets with other terms
Non-commercial clause (e.g., CC BY-NC)Exclude from commercial trainingMirrors that dropped the NC clause
Share-alike clauseCounsel review of how it applies to weights and derived dataDerived datasets you redistribute
No license foundExclude, or ask the original lab in writingAssuming "public" means permitted
Faces or bystanders visible, no consent recordExclude, or de-identify and get counsel sign-offBiometric statutes such as Texas 503.001 [5]
Facility or OEM terms unclearAsk the contributor; exclude meanwhileSite agreements behind lab recordings

Removing subsets changes the mixture, so rerun your embodiment and task coverage numbers afterward. The gap that remains is the brief for licensed data: specific robots, tasks, environments or failure cases you cannot cover with cleared open subsets.

Filling gaps with licensed operational data

When the cleared open mixture falls short, licensed recordings from companies that run robots or manual work can fill specific gaps. The decision between commissioning new collection and licensing existing recordings is covered in collection services versus licensing, and robot dataset acceptance checks covers what to verify before paying.

SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and manages licensing and ongoing purchases. Nothing is held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery; personal details are removed or replaced before delivery, with the method recorded and a sample checked, though no method is perfect. You can describe the robots, tasks and environments you need through the SourceX buyer request page. For the broader landscape, see the multimodal and embodied data hub, the robotics and embodied AI use case, the training data due diligence checklist and the glossary entry on dataset licensing.

Licensed robot data for commercial training

SourceX helps AI teams wherever they are based find US companies that hold the operational and hands-on work data they describe, then assesses data and licensing permissions before any agreement. Every release is approved by the supplying company, and nothing is contracted until a supplier agrees. Describe the robot data you need.

Frequently asked questions

Does a CC BY 4.0 statement on a pooled robot repo cover every subset?

Not by itself. An umbrella statement on a mixture's repository does not restate each contributor's terms, and pooled data such as Open X-Embodiment comes from many contributors [1], so confirm each subset's upstream terms and record any differences.

Can we keep non-commercial subsets for research and drop them later?

Some teams train internal research checkpoints on a wider mixture, but weights trained on NC data can carry that restriction into later use. Keep separate mixture configs and model lineage so a commercial model never inherits an NC subset.

Is converting RLDS episodes to another format a new license?

Generally not. Reformatting episodes, re-chunking trajectories or renaming fields like languageinstruction leaves the original terms attached to the data, and a converter's code license or hub tag does not replace them [3][4]. Ask counsel whether your converted copy counts as an adaptation under the upstream license.

Sources

  1. Open X-Embodiment Collaboration (arXiv), "Open X-Embodiment: Robotic Learning Datasets and RT-X Models" (2023). https://arxiv.org/abs/2310.08864v1
  2. Khazatsky, Pertsch et al. (arXiv), "DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset" (2024). https://arxiv.org/pdf/2403.12945
  3. Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  4. Hugging Face, "Create a dataset card (Datasets library documentation)". https://huggingface.co/docs/datasets/v2.19.0/en/dataset_card
  5. Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier" (2026). https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  6. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  7. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  8. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data