Multimodal and embodied data
Camera, LiDAR and Radar Sensor-Fusion Data: Sourcing and Specification
Quick answer
A usable sensor fusion dataset is not a folder of synchronized video and point clouds. It is calibrated, time-aligned camera, LiDAR, radar and IMU/GNSS recordings with per-sensor intrinsics, extrinsics, ego pose and a documented sync method, plus 3D labels that match your task. Public sets such as Argoverse 2 show the shape of the deliverable, but many restrict commercial use. Specify calibration, sync, label types and long-tail coverage first, then verify the license on every source.
By SourceX Editorial · Updated
What a sensor-fusion dataset must contain to train perception models
A fusion dataset must let you project any LiDAR point or radar return into any camera image at the correct instant, or it is not a fusion dataset. That requirement comes from the models: production vehicles and robots combine cameras with range sensors, so camera-only benchmarks cannot train or evaluate fusion detection and tracking. Public multi-sensor releases built for perception and forecasting, such as Argoverse 2, package ring cameras, LiDAR sweeps, poses and 3D annotations together for this reason [1].
The minimum content per recording session is:
- Raw sensor streams with native timestamps: images (often JPEG or raw Bayer), LiDAR sweeps (x, y, z, intensity, ring, per-point time), radar detections or tensors, IMU at high rate and GNSS fixes.
- Calibration: camera intrinsics and distortion model (pinhole with radial-tangential, or fisheye/equidistant), and a 4x4 extrinsic transform from each sensor to a common vehicle or robot frame.
- Ego pose: a 6-DoF trajectory in a world or map frame, with the method (RTK GNSS, INS fusion, LiDAR SLAM) stated.
- Labels in that same frame, with the annotation spec and class taxonomy.
- Sensor models and mounting: make, model, firmware, mounting height and orientation, and any changes during collection.
Containers vary. ROS bags and MCAP keep raw message streams; nuScenes-style relational JSON keeps keyframes and annotations; Parquet tables suit large label and pose sets. Our guide to robotics and sensor-fusion data formats covers the tradeoffs.
Calibration and time synchronization are where fusion data fails
Fusion data more often becomes unusable through calibration drift or timestamp misalignment than through missing frames. A LiDAR spinning at 10 Hz sweeps its field of view over about 100 ms, so an unknown offset between camera exposure and LiDAR azimuth smears objects when projected. Real rigs rarely run every sensor at one rate: a 10 Hz LiDAR, a camera near 12 to 30 Hz and an INS pose stream at 100 Hz or more is a typical mismatch, and the dataset must say how frames were associated (nearest timestamp, interpolation or hardware trigger).
Ask the supplier which clock each sensor stamps against (hardware trigger, PTP/IEEE 1588, GNSS PPS, or host receive time) and whether LiDAR points carry per-point timestamps for motion compensation. Host receive time is the weakest option, because USB and network buffering add variable latency.
Extrinsics drift when vehicles are serviced, sensors are bumped or brackets flex with temperature. Ask for the calibration procedure (target-based or targetless), the date of each calibration and a validation metric such as reprojection error. Radar extrinsics are the hardest to verify because radar returns are sparse and noisy, so treat a single factory calibration covering months of fleet data with caution and ask for per-session checks.
4D radar data needs its own specification
Imaging or 4D radar adds elevation to range, azimuth and Doppler, and its value depends on what level of the signal chain you receive. Research datasets increasingly publish this 3+1D form alongside LiDAR and stereo cameras, which lets teams use LiDAR boxes to supervise radar detectors. Most commercial recordings, however, expose only the radar's post-processed point list, not the raw ADC data or range-Doppler cube.
Specify which representation you need:
- Point list (detections): x, y, z or range/azimuth/elevation, radial velocity, RCS, SNR. Compact and available from most automotive radars.
- Radar tensor or cube: range-azimuth-Doppler bins before CFAR thresholding. Large, rarely retained, and often vendor-locked.
- Ego-motion-compensated Doppler: radial velocity with ego speed removed, which requires accurate pose at radar time.
Also ask whether multiple radar scans were accumulated before labeling, since accumulation changes point density and the comparability of evaluation results.
3D label types and annotation specs
Label type determines which models the data can train, so name it before you discuss volume. The common options are:
| Label type | Trains | Spec details to require |
|---|---|---|
| 3D bounding boxes | 3D detection | Box parameterization (center, size, yaw), coordinate frame, visibility and occlusion attributes, minimum points per box |
| Tracks (box + persistent ID) | Multi-object tracking, prediction | ID persistence across occlusion, keyframe vs interpolated frames, track start/end rules |
| Point-level semantic or panoptic segmentation | LiDAR segmentation | Class taxonomy, handling of unlabeled and noise points, ground vs drivable surface |
| Occupancy (voxel grids) | Occupancy networks | Voxel size, extent, how occlusion and free space are derived, whether generated from aggregated LiDAR |
| 2D boxes or masks on images | Camera branch, projection checks | Whether 2D and 3D labels share object IDs |
Ask how labels were produced (manual, model-assisted with human review, or auto-labeled offline) and request inter-annotator agreement or audit rates per class. Auto-labeled occupancy built from aggregated LiDAR inherits every pose and calibration error underneath it.
Coverage beyond road driving: warehouse, agriculture and mining
Off-road and indoor autonomy rarely matches public road datasets in sensor layout, classes or failure modes. A warehouse AMR may run a low-mounted 2D or solid-state LiDAR with fisheye cameras; an agricultural platform faces dust, crop canopy occlusion and rows that confuse ground segmentation; mining haul trucks see airborne particulate that produces false LiDAR returns and radar multipath off rock walls.
Long-tail scenarios drive value in every domain: night, rain, fog, dust, snow, sensor glare, rare objects and near-miss interactions. Ask the supplier for a scenario distribution table by condition and class, not an hour count. For operational logs from fleets already working in these environments, see operational logs from deployed robot fleets.
Licensing open and commercial point cloud data
Treat every public sensor fusion dataset as non-commercial until you have read its own license page. As of October 2026, the Waymo Open Dataset ships under its own license agreement with non-commercial restrictions [2], and other widely used sets such as Argoverse 2 [1], nuScenes and KITTI publish their own terms that you should confirm before production training. A large audit of popular text fine-tuning datasets found license metadata on aggregators is often missing or wrong, and the same risk applies to perception data, so trace each set to its original license text [4]. Mirrors and derived versions on model hubs frequently drop the original terms, and a non-commercial clause on the source generally still applies to redistributed and derived copies.
For private fleet data, ownership is often split among the vehicle or robot operator, the sensor or platform OEM and the site owner whose facility appears in the scans. Our page on who owns robot data walks through that split. If a dataset combines recordings from several operators, also check licensing multimodal records assembled from several rightsholders.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Privacy: faces, plates and locations
Camera streams in fusion data carry faces and license plates, so they need anonymization before licensing. Some state laws treat a record of face geometry as a biometric identifier requiring notice and consent before commercial capture, as Texas does [5]. Anonymization has a training cost: one study found traditional methods such as blurring degraded models trained on anonymized data, most when full bodies were anonymized, while realistic anonymization reduced the loss and left a minimal drop for face anonymization [3].
LiDAR rarely identifies a person directly, but dense point clouds and ego poses can reveal private sites, facility layouts and home addresses at trip endpoints. Ask whether GNSS traces were truncated near origins and destinations, and whether map-frame coordinates were offset. Our guide to de-identifying multimodal records covers face, screen and metadata handling together, and location-heavy data overlaps with geospatial and location data.
Sensor-fusion data request specification
A written specification is the fastest way to get comparable answers from suppliers. Adapt the template below, and see the broader multimodal dataset specification template for fields shared with other modalities.
Illustrative example: invented to show structure; it does not describe an available dataset.
request: sensor_fusion_perception_v1
domain: warehouse_amr # road | warehouse | agriculture | mining
sensors:
cameras: {count: 4, model_known: required, intrinsics: required, distortion_model: required}
lidar: {type: solid_state, per_point_timestamp: required, fields: [x, y, z, intensity]}
radar: {type: 4d_imaging, representation: point_list, doppler: ego_compensated}
imu_gnss: {imu_rate_hz_min: 100, gnss: optional_indoor}
calibration:
extrinsics: 4x4_to_base_link
method: documented # target-based or targetless
recalibration_log: required
validation: reprojection_error_px
time_sync:
clock_source: ptp_or_hw_trigger # host receive time not accepted
max_cross_sensor_offset_ms: 5
ego_pose: {frame: map, method: documented, rate_hz_min: 50}
labels:
types: [3d_boxes, tracks]
taxonomy: attach_class_list
qa: {audit_rate_per_class: required}
coverage:
conditions: [low_light, dust, wet_floor, reflective_racking]
report: scenario_distribution_table
privacy:
faces: anonymized_realistic_preferred
plates: anonymized
location: endpoints_truncated
delivery:
container: mcap_or_parquet_plus_json_calibration
sample: one_full_session_with_calibration
license:
uses: [training, evaluation]
confirm: commercial_use, derivative_models, term
When the data you need sits in private fleets rather than public releases, you can send this specification to SourceX as a buyer request. Score each supplier response against the same fields, and reject samples that arrive without calibration files, because you cannot validate projection quality from imagery alone. Telemetry-only needs without perception labels fit better under sensor and IoT data.
Sourcing sensor fusion datasets through SourceX
SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and manages the licensing process; nothing is held in stock and a request does not guarantee a match. You describe the data, every release is approved by the supplying company, and each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery. Start from the Multimodal and embodied data hub or describe your sensor-fusion data need to SourceX.
Frequently asked questions
Can I build a commercial model on nuScenes, Waymo Open or Argoverse 2?
Check each dataset's own license page rather than assuming. The Waymo Open Dataset uses a custom agreement with non-commercial restrictions [2], and other public fusion sets, including Argoverse 2 [1] and nuScenes, publish their own terms that you should read before any commercial training.
How do I verify calibration quality in a sample?
Project LiDAR points onto every camera image across a full session, including turns and hard braking. Edge misalignment that worsens during motion points to time-sync problems; constant misalignment points to extrinsic error.
Is radar point-list data enough for radar-camera fusion?
For most detection and tracking models, yes, if Doppler and RCS are included and timestamps are aligned. Research on raw radar representations needs tensors that most fleets do not retain.
Sources
- arXiv (Wilson et al.), "Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting" (2023). https://arxiv.org/pdf/2301.00493
- Waymo, "Waymo Open Dataset License Agreement (terms)". https://waymo.com/intl/es/open/terms/
- CVPR Workshops (Hukkelas and Lindseth), "Does Image Anonymization Impact Computer Vision Training?" (2023). https://openaccess.thecvf.com/content/CVPR2023W/WAD/papers/Hukkelas_Does_Image_Anonymization_Impact_Computer_Vision_Training_CVPRW_2023_paper.pdf
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/pdf/2310.16787.pdf
- Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.