Multimodal and embodied data
Tactile and Force-Torque Data for Contact-Rich Robot Learning
Quick answer
Tactile and force-torque data for robot manipulation means contact signals, either optical tactile images, taxel pressure arrays or 6-axis wrench streams, recorded in sync with camera frames, proprioception and commanded actions. Buyers should specify the exact sensor model, native sampling rates, clock synchronization method, calibration and bias records, and task outcome labels. Public sets such as Touch and Go and REASSEMBLE are useful references, but commercial training usually needs licensed recordings from your target sensor and tasks.
By SourceX Editorial · Updated
Why contact data is sourced separately from vision-only demonstrations
Contact data is a separate sourcing problem because most large robot corpora were built around cameras and joint states, not touch. Pooled collections such as Open X-Embodiment (over one million trajectories across 22 embodiments) [1] and DROID (76k trajectories, 350 hours across 564 scenes) [2] are strong for visual generalization, but buyers should not assume they carry wrist wrench or fingertip tactile channels for each episode. Check the per-episode schema before you plan training around them.
The tasks that need touch are the ones where vision is occluded or ambiguous at the moment that matters: peg-in-hole and connector insertion, wiping with controlled normal force, cable routing, and grasping deformables such as bags, fabric or food. For those tasks, purpose-built research sets exist, such as visuo-tactile collections built with GelSight or DIGIT sensors and assembly datasets that add wrist force-torque and audio to multi-view video. They are typically far smaller than vision-only corpora and recorded on a single lab setup.
That scale gap is the commercial opportunity and the risk. Contact datasets are small, sensor-specific and often research-licensed, so most teams end up commissioning or licensing recordings. The broader multimodal training data hub covers how this fits with other embodied data purchases.
Tactile image, taxel array or wrist wrench: choosing the modality
The modality you buy should match the sensor your policy will run on, because tactile data transfers poorly across sensor families. Vision-based tactile sensors such as GelSight, DIGIT and GelSlim image the deformation of an illuminated elastomer with an internal camera. Their output is a video stream of gel images, and gel stiffness, marker patterns, lighting and camera resolution all differ between models and even between gel batches.
Taxel arrays (capacitive, piezoresistive or magnetic skins on fingertips or palms) output a low-resolution pressure or 3-axis force grid per pad. A wrist-mounted 6-axis force-torque sensor outputs Fx, Fy, Fz, Tx, Ty and Tz at the flange, which captures net interaction forces but not where on the fingers contact happened. Joint torque estimates from collaborative arm controllers are a cheaper proxy, though they mix in friction and model error.
Public visuo-tactile sets such as Touch and Go, ObjectFolder, SSVTP and YCB-Slide come from different sensor designs, so they do not pool cleanly. In practice a policy pretrained on GelSight images will usually need fine-tuning or adaptation for DIGIT, and taxel data needs its own encoder. Name the sensor model and firmware in your request, and treat cross-sensor data as pretraining material, not a drop-in replacement.
| Modality | Typical output | Best for | Main failure mode to check |
|---|---|---|---|
| Vision-based tactile (GelSight, DIGIT, GelSlim) | RGB gel images per fingertip | Slip detection, texture, in-hand pose, edge localization | Gel wear, lighting drift, unrecorded gel or sensor swaps |
| Taxel array skin | Pressure or 3-axis force grid per pad | Grasp force distribution, whole-hand contact | Dead taxels, crosstalk, hysteresis |
| Wrist 6-axis F/T | Fx, Fy, Fz, Tx, Ty, Tz at the flange | Insertion, wiping, compliance control | Bias drift, temperature effects, tool mass not compensated |
| Controller joint torque | Per-joint torque estimates | Low-cost contact detection on cobots | Friction and model error masking small contacts |
Synchronization, rates and calibration: the specification that decides usability
A contact dataset is only usable if every stream shares a recorded timebase and every sensor's calibration state is known. Force-torque sensors typically sample much faster than cameras, so ask suppliers to deliver the native-rate stream plus the resampled training view, not only a downsampled copy aligned to video frames. Tactile cameras have their own frame clock and exposure latency, which should be logged.
Ask how clocks were aligned: hardware trigger, PTP, or software timestamps on a single host. ROS 2 bags or MCAP files usually carry both header stamps and receive times; request both, because the difference exposes transport latency. Without this, even a small offset between a wrench spike and the camera frame can teach a policy to react after the event it should anticipate.
Calibration is the other half. Wrist sensors need a recorded bias (tare) procedure, the tool and gripper mass used for gravity compensation, and the sensor-to-flange transform. Strain-gauge sensors drift with temperature and time, so ask for per-episode zero offsets and ambient temperature where available. For gel sensors, ask for reference no-contact frames per session and a log of gel replacements, because a new gel changes the image distribution.
Industrial records that already contain force signals
Some of the most useful contact data already sits in factory and test systems, not robotics labs. Robot controllers with force-control options log contact forces during guided insertion and polishing. Electric nutrunners and DC tightening tools record torque-angle curves for every fastener, with pass or fail results against torque and angle windows. Servo press-fit stations record force-displacement curves with envelope or window evaluations.
These records are attractive because they are large, outcome-labeled and tied to real part variation. Their limits are equally clear: they rarely include synchronized vision, the "action" is a fixed process program rather than a learned policy, and units, sampling and evaluation windows differ between tool vendors. They fit force-profile models, anomaly detection on insertions and reward or success classifiers better than end-to-end visuomotor policies. The sensor and IoT data buying guide covers general telemetry; for assembly sequences with tool signals and work instructions, see assembly demonstration data with tool signals.
When you request this kind of data, describe the process, tool type and curve format you need. Ownership questions between the operating plant, integrator and tool vendor are covered in who owns robot data.
A request specification for contact-rich data
A good request names the sensors, the tasks and the acceptance tests before any price discussion. The template below is a starting point; adapt field names to your training stack, for example LeRobot or RLDS episode schemas.
Illustrative example: invented to show structure; it does not describe an available dataset.
request: contact_rich_insertion_v1
tasks:
- connector insertion (USB-C, RJ45), peg-in-hole with chamfer variants
- wiping a curved panel at controlled normal force
robot: 7-DoF arm, parallel gripper; record make, model, controller version
sensors:
wrist_ft: {type: 6-axis strain gauge, channels: [Fx,Fy,Fz,Tx,Ty,Tz], units: [N, Nm], native_rate_required: true}
tactile: {type: vision-based gel, model: named, per_finger: true, format: mp4 lossless or png frames}
cameras: [wrist RGB, two static RGB-D]
proprio: [joint_pos, joint_vel, joint_torque_est, ee_pose, gripper_width]
actions: commanded ee_delta or joint targets, plus controller mode (impedance/position)
sync: hardware trigger or PTP; keep header and receive stamps; report measured offsets
calibration: per-episode F/T bias, tool mass and CoM, sensor-to-flange transform, gel no-contact frames
labels: success/failure, failure type (jam, misalignment, slip, drop), operator interventions
format: MCAP or ROS 2 bag raw; parquet or RLDS training view; JSON Lines manifest
volume: pilot slice first; full set scoped after acceptance checks
Run your own acceptance tests on a sample before committing: verify channel completeness per episode, check that wrench spikes align with visible contact frames, confirm bias offsets are within your tolerance, and check force and tactile channels for clipping or saturation. Our acceptance checks for robot datasets page lists the general tests; add contact-specific ones such as saturation counts and drift over an episode.
Licensing and provenance risks specific to contact data
Contact datasets carry the same license ambiguity as other open data, plus hardware-specific gaps. An audit of more than 1,800 text datasets found license omissions above 70% and error rates above 50% on popular hosting sites [3], so confirm the license on the project page and paper, not an aggregator listing. Several research sets are released under non-commercial or unstated terms, and some papers promise a future data release that never lands on a stable download page.
Recordings made with human hands, such as human-collected tactile probing paired with egocentric video, can capture faces, voices or workspaces in frame. Check consent records and redaction, as covered in de-identifying multimodal records. For commercial use, also confirm who owns robot-generated logs at the facility where they were recorded and whether the tool or sensor vendor's software terms restrict export. For open sets, see checking open robot dataset licenses.
Where SourceX fits for tactile and force data
SourceX's fit for this category is narrow, and buyers should know that up front. SourceX sources operational datasets from US companies, including engineering records and new recordings of hands-on work, and manages the licensing process. Data is sourced on request rather than held in stock, so a request for force or tactile recordings does not guarantee a match.
Where a US business holds relevant records, such as process force curves or recorded manipulation work, every release is approved by that company and each dataset is rights-reviewed for ownership and consents before delivery under a license defining records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe the data you need, not the businesses, on the SourceX buyer page. For a broader view of embodied data, see robotics training data for embodied AI.
Request tactile and force-torque data for your manipulation models
SourceX looks for US businesses that hold the data you describe and runs the process from Find and Assess through Agree, Transact and Manage, with nothing contracted until a supplier agrees. Terms and pricing are agreed per deal. Describe your sensors, tasks and acceptance tests at sourcex.si/buyers.
Frequently asked questions
Can I train on GelSight data and deploy on DIGIT sensors?
You can pretrain on one and adapt to the other, but expect a distribution shift. Both are vision-based gel sensors, yet gel optics, lighting and resolution differ, so budget for fine-tuning data from the deployed sensor.
Is wrist force-torque enough without fingertip tactile sensing?
For many insertion and wiping tasks, a wrist 6-axis sensor carries most of the useful contact signal, and several research assembly datasets rely on it. Slip detection, in-hand pose and deformable grasping usually benefit from fingertip tactile signals.
Should I buy simulated contact data instead?
Simulated contact is useful for pretraining, but friction, gel deformation and sensor noise are hard to model faithfully. Our comparison of real vs simulated robot data covers when physical recordings justify the cost.
Sources
- arXiv (Open X-Embodiment Collaboration), "Open X-Embodiment: Robotic Learning Datasets and RT-X Models" (2023). https://arxiv.org/abs/2310.08864v1
- arXiv, "DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset" (2024). https://arxiv.org/abs/2403.12945v2
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.