Short answer
Humanoid robot training data is the mix of robot trajectories, human demonstrations, language, object, failure and evaluation data used to build and test a humanoid's models. Control-grade data, such as teleoperated trajectories, motion capture and force or tactile signals, comes from robots and specialist capture. Business assets such as workplace video, SOPs, product catalogs and inspection records feed perception, task understanding, planning and evaluation.
Why this is several purchases, not one
A humanoid model is trained in layers, and each layer eats a different kind of data. Recent foundation-model reports make the structure explicit. GR00T N1 describes a data pyramid: web data and human egocentric video at the base, synthetic and simulated trajectories in the middle, and real-robot teleoperation at the top, with latent actions and inverse-dynamics pseudo-actions used to label video that has no actions (NVIDIA, 2025).
The top of that pyramid is pooled across many labs. Open X-Embodiment gathered more than one million robot trajectories from 22 embodiments contributed by 21 institutions (Open X-Embodiment Collaboration, 2023). Even DROID, one of the larger single-rig efforts, reports about 76,000 teleoperated trajectories and roughly 350 hours across 564 scenes, recorded at 15 Hz with per-episode camera calibration (Khazatsky et al., 2024).
So a request for "humanoid training data" usually hides five or six separate purchases. Each has its own suppliers, preparation work, rights questions and failure modes. This section of SourceX maps those purchases to the places they can realistically come from.
The layers of a humanoid data program
| Layer | What it trains or tests | Typical assets | Where it comes from |
|---|---|---|---|
| Robot action data | Low-level control from observations to joint or end-effector commands | Teleoperated episodes, robot logs, handheld-gripper demonstrations | Robots in labs or deployments; specialist collection |
| Human demonstrations | Visual representations, task and step structure, latent actions; co-training when hands are tracked in 3D | Workplace task video, egocentric and wrist capture, 3D hand pose | Existing business footage; commissioned recording |
| Language and procedures | Task decomposition, instruction grounding, preconditions and success checks | SOPs, work instructions, checklists, narrated video | Business documentation and training teams |
| Objects | Detection, identification, pose, grasp planning, simulation assets | Catalog images, attributes, dimensions, CAD and scans | Retailers, distributors, manufacturers; commissioned imaging |
| Failures and quality | Defect, error and failure recognition; success detection | Inspection archives, quality records, mistake video, robot failure logs | Manufacturers and operators; robot operators |
| Evaluation | Held-out measurement of perception, planning or policies | Unseen sites, operators and objects; physical object sets | Any route, kept apart from training |
Requirements at each layer depend on architecture, task, embodiment and intended use. Published figures describe what one study used, not a floor. Gemini Robotics, for example, reports learning some new short-horizon tasks from as few as 100 demonstrations, but only on top of a large pretrained model (Google DeepMind, 2025).
Five distinctions that decide what to buy
1. Human workplace video is not robot action data. Video of people working carries no joint states, commanded actions or torques. It can pretrain visual encoders, teach what tasks and steps look like, supply latent actions or progress rewards, and, when hands are tracked in 3D, be co-trained with robot data; co-training studies report the human data as improving robot-only training rather than replacing it (Kareer et al., 2024; Qiu et al., 2025). The full account is in what human demonstration video can and cannot do for robot learning, and the boundary with robot data is drawn in human demonstrations vs teleoperated robot demonstrations.
2. Procedures are not physical demonstrations. SOPs and work instructions give ordering, constraints, tools and success criteria in language, but no motion, force or perception data. Language-grounded planning work divides the labor this way: the language model supplies task knowledge, and the robot's pretrained skills decide what is feasible (Ichter et al., 2022). See what SOPs can and cannot teach a humanoid.
3. Recognizing an object is not manipulating it. Catalog images, attributes and 3D models support detection, identification, pose estimation and grasp planning; ABO, for instance, pairs product catalog images and metadata with artist-made 3D models (Collins et al., 2021). None of that records how an item slips, flexes or swings once it is in a hand. See object recognition data vs manipulation training data.
4. Existing footage is not purpose-built collection. Archives were recorded to train staff, monitor a process or settle disputes, so viewpoints, coverage, labels and consent were set for that purpose. A purpose-built corpus is designed from the start: Ego4D recruited consenting camera wearers and applied de-identification across 931 wearers and 74 locations (Grauman et al., 2021). Commissioned collection buys that control at the cost of time, money and site access; see existing workplace footage vs commissioned recording.
5. Training data is not evaluation data. An evaluation set must be held out by site, operator, object or time, whichever the claim requires, and never mixed back into training. Buy or build it as a separate asset with its own split rules, as set out in training datasets vs evaluation datasets.
Taken together: business footage and SOPs, on their own, will not train a humanoid's physical control system. They can make robot data go further.
Three sourcing routes
Existing business assets
Operating businesses already hold training-video libraries, process recordings, SOPs, LMS courses, product catalogs, machine-vision archives, work orders and WMS or MES event logs. These assets show real work at real sites and often exist in volume. Their weaknesses come from their origin: camera placement, labeling and consent were chosen for another job, and some rights sit with third parties such as course vendors, 3PL clients or equipment makers.
SourceX looks for US businesses that hold the data a buyer describes, checks the data and the supplier's licensing permissions, and agrees pricing and allowed uses in a license. Nothing is contracted until a supplier agrees.
Commissioned collection
New recording lets you choose viewpoints (head, wrist, fixed multi-view), sensors such as depth or IMUs, task scripts taken from real SOPs, operator and site diversity, deliberate failures and recoveries, and consent written for AI training. It still needs willing businesses, briefed demonstrators, calibration and quality control. The site describes a separate physical-world path for new recordings of hands-on work coordinated with businesses, where each opportunity has its own requirements. If no existing asset fits, SourceX can scope a commissioned collection of new recordings.
Specialist partners for robot-specific data
Robot-specific observations and actions come from robots, teleoperation rigs, handheld capture devices and instrumented setups, not from ordinary business archives; a business that already runs robots may hold logs, but the robot vendor often controls them. Buyers must specify the embodiment, action space (joint or end-effector, absolute or delta), control frequency, calibration and sensor synchronization. Handheld grippers sit between human video and teleoperation: the Universal Manipulation Interface records demonstrations with a wrist-mounted fisheye camera as its only sensor (Chi et al., 2024). Contact-rich datasets add force: RH20T pairs each robot sequence with visual, force, audio and action data plus a human demonstration video (Fang et al., 2023).
The specialist pages cover teleoperation data collection requirements, robot trajectory dataset specification, motion capture for humanoid motion, tactile sensing data and force/torque data for contact-rich tasks. SourceX's role is sourcing, rights review, licensing and delivery coordination; it can also arrange annotation and engage specialist capture partners, scoped case by case.
To choose between the three, use existing assets, new collection or specialist partners.
The eight clusters
| Cluster hub | Start here if you need | Limit to plan for |
|---|---|---|
| Workplace task demonstrations | Video of real work for pretraining, task understanding and co-training; capture and annotation specs | No robot actions or forces |
| SOPs and procedural data | Task structure, instruction grounding, planner context, success checks | No motion or perception; written practice can drift from actual practice |
| Product and object data | Recognition, attributes, dimensions, 3D models, novel-object test sets | Shows objects, not their behavior in the hand |
| Inspection, defect and failure data | Defect images, human error video, robot failure examples | Leakage across lot and line; many public sets are non-commercial |
| Maintenance, repair and tool use | Work orders, service documentation, repair and tool-use video | Equipment-maker documents are hard to license; no torque feel |
| Warehouse and fulfillment | Picking, packing and unloading footage plus system-generated labels | Client-owned goods appear in 3PL footage |
| Manufacturing and assembly | Assembly video, standard work, routings, time studies | Design and customer confidentiality; insertion forces |
| Procurement, evaluation and specialist collection | Specifications, licensing, consent, acceptance, robot-specific data | Diligence, not legal advice |
Public benchmarks are a reminder to check terms early. MVTec AD, a common inspection benchmark, is licensed CC BY-NC-SA 4.0 and excludes commercial use (Bergmann et al., 2021), which can push commercial teams toward licensed alternatives.
Keystone pages to read first
- Decide and specify: existing assets, new collection or specialist partners; requirements specification template; human demonstrations vs teleoperation; existing footage vs commissioned recording.
- Know the limits: human demonstration video for robot learning; the human-to-humanoid embodiment gap; what SOPs can and cannot teach; recognition vs manipulation data.
- Assess common business assets: corporate training video archives; instruction-paired demonstration datasets; product catalog data for robot perception; defect image datasets; work order histories; 3PL providers as data suppliers.
- Rights: licensing terms for robotics training data; worker consent and identifiable people.
Illustrative: one program mapped to routes
A synthetic example, program_A, targets tote handling and shelf restocking in store backrooms.
| Need | Asset to look for | Route | Still missing |
|---|---|---|---|
| Visual pretraining on real backroom scenes | Store or 3PL handling footage, site_003 to site_011 | Existing assets | Hand detail; consent scope |
| Step structure and language | Store operating procedures and planograms | Existing assets | Mapping steps to robot skills |
| Item recognition | Catalog images and dimensions for target SKUs | Existing assets | In-situ appearance, clutter |
| Hand-tracked demonstrations | Head and wrist capture with 3D hand pose, operator_A01 to operator_A24 | Commissioned collection | Robot actions |
| Control policy data | Teleoperated episodes on the target humanoid | Specialist | Business context at scale |
| Held-out evaluation | Two stores and 40 SKUs never used in training | Commissioned or existing, held out | Nothing, if the split is enforced |
No row stands in for another. The commissioned video narrows what the teleoperation must cover; it does not remove it. GR00T N1's authors, for instance, used generated neural trajectories to expand 88 hours of teleoperation data to 827 hours, but the real teleoperation remained the anchor (NVIDIA, 2025).
How to judge any offer
- Layer. Which part of your stack does this asset feed, and what will still come from robot data?
- Origin. What was it recorded or written for, and does that purpose limit viewpoint, coverage or consent?
- Gaps. Are actions, forces, calibration, step labels or language missing, and can any be added afterwards?
- Coverage. Counts by site, operator, object and condition, including failures, not only totals.
- Rights. Who can license it, and does every person's consent and every third party's content allow AI training?
- Separation. Can part of it be held out cleanly for evaluation?
- Evidence. A sample you can inspect before signing, and acceptance tests before payment.
For detail, use the robotics dataset evaluation checklist, provenance and chain of custody and acceptance testing for dataset deliveries.
Preparing a request
The SourceX buyers page leads to a request form that asks for dataset kind (including "Physical-world collection"), modality, approximate volume, format, geography, language, timeframe, licensing requirements, de-identification requirements and evaluation criteria. The requirements specification template shows how to fill those fields for robotics data, including the parts a supplier needs most: target tasks, viewpoints, required labels and what must be held out.
Discuss your requirements
Bring the layer you are trying to fill, the tasks and environments, and the uses you need licensed. SourceX reviews the requirement, investigates whether US businesses hold suitable material, and follows up about potential sources; a request does not guarantee a match. Buyers can be anywhere. For broader context, the catalog covers robotics and embodied AI use cases and first-person video of skilled manual work.
Discuss your humanoid training-data requirements with SourceX
Sources
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots — data pyramid, pseudo-actions for action-less video, teleoperation expanded with neural trajectories.
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models — robot data pooled across embodiments and institutions.
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset — scale and capture parameters of a large teleoperated dataset.
- Gemini Robotics: Bringing AI into the Physical World — few-demonstration adaptation on top of a pretrained model.
- EgoMimic: Scaling Imitation Learning via Egocentric Video — co-training tracked human and robot data.
- Humanoid Policy ~ Human Policy — egocentric human demonstrations as cross-embodiment data for humanoids.
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances — language for task knowledge, robot skills for feasibility.
- ABO: Dataset and Benchmarks for Real-World 3D Object Understanding — catalog images and 3D models for object understanding.
- Ego4D: Around the World in 3,000 Hours of Egocentric Video — design of a consented, purpose-built egocentric corpus.
- The MVTec Anomaly Detection Dataset — non-commercial license on a common inspection benchmark.
- Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots — handheld-gripper demonstration capture.
- RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot — robot sequences with force, audio and action data.