Skip to content

Video data

Video Data for World Models: Physical Interactions, Camera Motion and Coverage

Quick answer

World model training data is video whose value comes from what happens between frames: objects colliding, deforming and flowing, cameras moving through 3D space, and, ideally, the actions that caused each change. Buyers should specify action conditioning (driving controls, equipment commands, device inputs), camera pose or the means to estimate it, long uncut clips with causal sequences, and deliberate coverage of contact-rich physics. Aesthetic quality and captions matter far less than they do for text-to-video.

By SourceX Editorial · Updated

Why world-model video is judged differently from generation video

A world model is evaluated on whether its predicted futures stay physically consistent, so the buyer cares about dynamics, not looks. The page on text-to-video training data covers aesthetics, captions and prompt alignment; here the questions are whether objects persist, whether contact produces plausible responses and whether the camera's own motion is separable from scene motion. Expect a world-model curation pipeline to cut source video at scene changes, keep dynamic, information-rich clips and de-duplicate near-identical footage before training.

That pipeline tells you what large teams discard. Static talking heads, slideshows, heavy editing and repeated near-duplicate footage add tokens without adding dynamics. When you license real footage, ask the supplier to run the same shot-boundary and motion filters before quoting volume, so you are not paying for hours your curator will throw away.

Passive video versus action-conditioned video

Action-conditioned video pairs each frame with the control input that produced the next state, and it is the scarcer and more valuable half of a world-model corpus. Passive video teaches what tends to happen; action-conditioned video teaches what happens because of a specific input, which is what planning requires. Robotics corpora show what the paired form looks like: DROID records teleoperated manipulation with camera streams and robot actions together [3], and Open X-Embodiment pools many such datasets across robot types [4].

A common design is a large passive base with a smaller, tightly synchronized action slice on top, so buy each half against its own spec. Real sources of action streams include:

  • Vehicle telemetry: steering angle, throttle, brake pressure, gear and speed from CAN bus logs alongside dashcam or surround video. See dashcam video and fleet footage.
  • Equipment commands: PLC tags, CNC G-code execution logs, forklift joystick and mast inputs, crane controls.
  • Teleoperation: robot joint states and end-effector commands, as in DROID and the pooled Open X-Embodiment collection [3][4].
  • Device and game inputs: keyboard, mouse and controller events with frame-accurate timestamps.

If action labels are missing, an inverse dynamics model can recover them. OpenAI's VPT trained an inverse dynamics model on a small contractor-labeled set and used it to pseudo-label large volumes of unlabeled video [1]. For buyers, that means a small, well-synchronized labeled slice from the same domain can raise the value of a larger passive archive.

Camera pose and ego-motion metadata

Camera pose separates "the world moved" from "the camera moved," and you should require it or confirm it can be estimated reliably. Driving datasets such as Argoverse 2 ship synchronized sensors with vehicle pose for exactly this reason [5]. For licensed footage, the useful fields are intrinsics (focal length, principal point, distortion model), extrinsics relative to a vehicle or body frame, and a per-frame 6-DoF pose from GNSS/INS, visual-inertial odometry or post-hoc structure-from-motion.

Common failure modes to screen for:

  • Rolling shutter and stabilization: rolling-shutter skew during fast motion and electronic image stabilization both warp frames and degrade geometric pose estimation; ask about the sensor type and whether raw or stabilized output was recorded.
  • Variable frame rate: phone and some dashcam encoders can fall back to variable frame rate under load, which corrupts action-to-frame alignment.
  • Clock drift: IMU, CAN and video clocks that are not hardware-synced can drift by milliseconds per minute; ask for the sync method (PTP, GPS PPS or post-hoc cross-correlation).
  • Missing intrinsics: without calibration, lens distortion and focal length must be guessed, which makes depth and pose estimates less accurate and harder to audit.

Long continuous clips and causal sequences

World models need clips long enough to contain cause, effect and recovery, so specify minimum uncut duration rather than total hours. Classic action-recognition sets like Kinetics use short trimmed clips per action class [7], which suit classification but rarely contain the full arc of a spill, a jam or a near-miss. Long-context work is now evaluated separately, as LongVideoBench does with videos up to an hour long [6].

Operational footage is a natural source of long causal sequences because cameras run continuously through shifts. Egocentric collections show the shape of this: Ego4D covers 3,670 hours of daily-life video from 931 camera wearers in 74 locations, collected with consenting participants [2]. Hands-on work recorded in manufacturing assembly or warehouse operations gives repeated tasks with natural variation, plus failures and corrections that scripted capture misses.

Coverage of contact, deformation and fluids

Physical interactions that web video under-represents are where licensed real footage earns its price, so write coverage targets into the request. Consumer web video skews toward people, faces, scenery and edited content; it is thin on sustained close-range contact, material failure and industrial processes. A world model that has never seen cardboard crush, cable bend or liquid pour is likely to produce implausible versions of those dynamics.

Useful coverage axes include rigid contact (stacking, collisions, tool strikes), deformables (cloth, cable, food, packaging), fluids and granular media (pouring, spraying, conveyors of bulk material), articulated objects (doors, drawers, valves), and occlusion with re-emergence. Add lighting and weather strata, viewpoint (egocentric, fixed overhead, vehicle-mounted) and failure rate. Simulation can fill rare cases; see synthetic versus real video data and real versus simulated robot data for where each holds up.

A request spec for world-model video

The fastest way to get comparable quotes is a structured spec that names modalities, sync tolerances and coverage before anyone discusses volume. The record below shows the level of detail suppliers need.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy it matters
clip_idveh-0412_2026-03-18T14:02:11Z_seg07Stable key across video, telemetry and labels
container / codecMP4, H.264 or H.265, constant 30 fpsVariable frame rate breaks action alignment
min_uncut_duration120 s, no edits or overlaysKeeps cause, effect and recovery in one clip
camera_intrinsicsfx, fy, cx, cy, distortion model per cameraNeeded for metric depth and pose
pose_stream6-DoF at 100 Hz, source GNSS/INSSeparates ego-motion from scene motion
action_streamsteering_deg, throttle_pct, brake_bar at 50 HzEnables action-conditioned training
sync_method / toleranceGPS PPS, within 10 msStates how alignment can be audited
interaction_tagscontact, deformable, fluid, occlusionMeasures physics coverage
redactionfaces and plates blurred, method loggedBiometric and privacy exposure
license_scopetraining and evaluation, internal modelsDefines allowed use before delivery

Ask for a sample of a few dozen clips with every stream attached, then check sync by finding visible events (a brake light, a tool strike) and confirming the action stream changes on the matching frame.

Rights, people and documentation in physical-world footage

Real-world footage carries layered rights, and world-model buyers should clear them before scale, not after. A single clip can involve the footage owner, recorded workers or bystanders, visible brands and ambient audio; the rights layers in a video clip page walks through each. Where workers are recorded, notice and consent rules apply; see recording employees on video.

Face geometry is a biometric identifier under Illinois BIPA, which sets retention and written-release duties [9]. On copyright, the U.S. Copyright Office's Part 3 report on generative AI training remained a pre-publication version as of October 2026 [10], so licensed provenance is the practical way to reduce uncertainty. Ask suppliers to document each dataset in a machine-readable format such as Croissant with its responsible-AI extension, covering collection method, consent basis and preparation steps [8].

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Where operational sources fit

Companies that already run cameras alongside machines or vehicles often hold the paired streams world-model teams want, but those streams sit in separate systems. Fleet operators keep dashcam video next to telematics; plants keep line cameras next to historians and PLC logs; logistics sites keep overhead video next to WMS scan events. Joining them is the work. See sensor and IoT data, robotics data for embodied AI and the physical world overview, and the human demonstration video for robot learning page for manipulation-specific needs.

SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and manages licensing and ongoing purchases; nothing is held in stock, and a request does not guarantee a match. It does not source scraped web content or generic CCTV. Teams can describe the video and action streams they need without naming businesses. For background, see the video data hub, the AI data hub and the definition of real-world data.

Sourcing action-conditioned world model training data

SourceX looks for US businesses that hold the footage and control logs you describe, reviews ownership and consents, and delivers only under a license that defines records, uses, term and delivery once the supplying company approves. Personal details are removed or replaced before delivery, with the method recorded and a sample checked. Start a world-model data request.

Sources

  1. OpenAI (arXiv:2206.11795; NeurIPS 2022), "Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos" (2022). https://arxiv.org/pdf/2206.11795
  2. Grauman et al., Ego4D consortium (arXiv:2110.07058), "Ego4D: Around the World in 3,000 Hours of Egocentric Video" (2021). https://arxiv.org/abs/2110.07058v1
  3. Khazatsky et al. (arXiv:2403.12945), "DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset" (2024). https://arxiv.org/pdf/2403.12945
  4. Open X-Embodiment Collaboration (arXiv:2310.08864), "Open X-Embodiment: Robotic Learning Datasets and RT-X Models" (2023). https://arxiv.org/pdf/2310.08864
  5. Wilson et al. (arXiv:2301.00493), "Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting" (2023). https://arxiv.org/pdf/2301.00493
  6. Wu, Li, Chen, Li (arXiv:2407.15754; NeurIPS 2024), "LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding" (2024). https://arxiv.org/abs/2407.15754v1
  7. Kay, Carreira, Zisserman et al., DeepMind (arXiv:1705.06950), "The Kinetics Human Action Video Dataset" (2017). https://arxiv.org/pdf/1705.06950
  8. Jain et al., MLCommons (arXiv:2407.16883), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  9. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  10. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data