Video data
Hand-Object Interaction Annotation in Video: Grasps, Contacts and Tool Use
Quick answer
A useful hand-object interaction dataset labels more than hands. For manipulation learning you need, per hand and per frame or segment: hand side and location, contact state, the active object (ideally as a mask), grasp type from a fixed taxonomy, the object's state change, and the tool-to-target relation when a tool acts on something else. Specify each layer, its sampling rate and its QA rule before collection, because these labels are hard to retrofit onto footage later.
By SourceX Editorial · Updated
This page is a label specification. For the footage itself and how it is captured, see the video data hub and the egocentric video of skilled manual work page; for verb-level time segments, see temporal action segmentation labels.
Which label layers a manipulation team actually needs
Most manipulation-oriented teams need six layers, and each one answers a different modeling question. Public benchmarks show that the layers stack: EPIC-KITCHENS VISOR, for example, pairs pixel-level masks for hands and active objects in egocentric kitchen video with hand-object relations [1]. A buyer spec should say which layers are mandatory and which are optional.
- Hand detection and side. Box or mask per hand, tagged left or right, with a persistent track ID across the clip. Without the side tag you cannot model bimanual role split.
- Contact state. At minimum no-contact, self-contact, other-person contact, portable-object contact and fixed-object contact [9]. Fixed versus portable matters for robotics, because pushing a drawer and lifting a cup need different policies.
- Active object. The object the hand is manipulating, as a box, polygon or mask, linked to the hand by ID. VISOR's choice of masks over boxes reflects that thin tools and partially occluded objects are poorly described by rectangles [1].
- Grasp type. A class from a published grasp taxonomy, such as the GRASP taxonomy of static, one-handed grasps organized by opposition type, power versus precision and thumb position [8], or a coarser merged version of it.
- Object state change. Pre-condition, post-condition and the frame where the change occurs (open to closed, whole to cut, loose to fastened).
- Tool-to-target relation. For tool use, a triple of hand, tool and target object with the effect on the target, for example screwdriver acting on screw, screw becoming seated.
Hand keypoints (2D joints) are a seventh layer that some teams add. Once you add depth, IMU or 3D hand pose and mesh, the dataset becomes a multimodal capture problem with calibration and synchronization requirements, which is a different specification from video labels alone.
How to pick a grasp taxonomy annotators can apply consistently
Pick the coarsest taxonomy your model can use, because inter-annotator agreement drops quickly as classes multiply. A full fine-grained grasp taxonomy is precise but hard to apply from a single camera angle where fingers are occluded. Many teams label a merged set of general grasp classes, or a three-way power, precision and intermediate split, and reserve fine classes for a calibrated subset.
Two rules prevent the most common confusion. First, label grasp only while contact state is portable-object contact and the grasp is stable, since grasp taxonomies describe static, stable grasps and transitional finger motion does not fit them. Second, add an explicit "non-prehensile" value for pushing, pressing and sliding, so annotators are not forced to pick the nearest grasp class for a hand that is not grasping.
Industrial work needs a domain vocabulary on top of the grasp classes. A general taxonomy will not tell you whether a hand holds a torque wrench, a crimper or a pipette; those come from a tool and part list agreed with domain reviewers before labeling starts. The manufacturing assembly video guide covers how assembly vocabularies are usually built.
How viewpoint changes occlusion, label cost and what the model learns
Egocentric and third-person video produce different labels for the same task, so specify viewpoint before you specify labels. Head-mounted footage, as in Ego4D's more than 3,000 hours of first-person video [2], keeps hands near the image center and large, which makes contact and active-object labels easier, but the wrist and the far hand often leave the frame. Fixed third-person cameras see both hands and the body, which helps bimanual role labels, but small tools shrink to a few pixels.
Multi-view capture resolves many occlusion disputes. Assembly101 recorded eight static and four egocentric views simultaneously of people taking apart and rebuilding toy vehicles [3], which lets an annotator check a hidden grasp in another view. If you buy multi-view footage, require frame-level synchronization metadata, or cross-view label propagation will drift.
Bimanual tasks need per-hand labels, not per-person labels. ATTACH annotated assembly actions by hand, including two-handed and simultaneous actions [4], because a single "person is screwing" label hides the stabilizing hand that a robot also has to reproduce. If your target platform is dual-arm, make per-hand labeling a hard requirement.
An illustrative label specification and record
The practical deliverable is a label specification that names every layer, its geometry, its sampling rate and its acceptance rule. Writing it down also forces agreement on what "contact" and "active object" mean before anyone draws a mask.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Layer | Geometry and values | Sampling | Acceptance check |
|---|---|---|---|
| Hand track | Mask or box, side L/R, track_id | Every annotated frame | Side swaps per track = 0 in audited clips |
| Contact state | none, self, person, portable_obj, fixed_obj | Every annotated frame | Agreement with second annotator on a gold subset |
| Active object | Mask (polygon or RLE), object_class, object_id | Frames with contact | Mask IoU against expert gold above an agreed threshold |
| Grasp type | Merged general grasp classes plus non_prehensile | Stable contact segments only | Confusion matrix reviewed per class |
| Hand keypoints | 21 2D joints, visibility flag per joint | Keyframes | Occluded joints flagged, not guessed |
| State change | pre_state, post_state, change_frame | Per interaction segment | change_frame within a tolerance window |
| Tool relation | tool_id, target_id, effect verb | Per tool-use segment | Target linked to an existing object_id |
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"clip_id": "clip_000412",
"frame": 1834,
"view": "egocentric_head",
"hands": [
{
"track_id": "h1",
"side": "right",
"bbox": [612, 388, 142, 120],
"contact_state": "portable_obj",
"active_object_id": "o7",
"grasp_type": "power_medium_wrap",
"keypoints_visible": 17
},
{
"track_id": "h2",
"side": "left",
"bbox": [301, 402, 130, 118],
"contact_state": "fixed_obj",
"active_object_id": "o2",
"grasp_type": "non_prehensile"
}
],
"objects": [
{"object_id": "o7", "class": "screwdriver", "segmentation": "rle:...", "role": "tool"},
{"object_id": "o2", "class": "panel_bracket", "segmentation": "rle:...", "role": "target"}
],
"interaction": {
"tool_id": "o7",
"target_id": "o2",
"effect": "fasten",
"pre_state": "screw_loose",
"post_state": "screw_seated",
"change_frame": 1902
}
}
Ask for an export your training stack already reads. COCO-style JSON keeps images, categories and annotations in separate sections and supports polygon or RLE masks and keypoint tasks [6], which covers hands, objects and joints; contact, grasp and relation fields then go in per-annotation attributes or a sidecar file keyed by annotation ID. Whatever the format, require a schema document and a versioned label vocabulary file with the delivery.
How to check hand-object labels before you accept a delivery
Treat hand-object labels as suspect until you have audited a sample, because even widely used benchmarks carry measurable label noise. An audit of 10 popular test sets estimated an average label error rate of at least 3.3% [7], and contact and grasp labels are more subjective than most image classes. VISOR's authors used a multi-stage annotation and checking process to keep masks consistent over time [1], which is the kind of process detail to ask a supplier to describe.
Illustrative example: invented to show structure; it does not describe an available dataset.
Acceptance checklist for a hand-object annotation delivery
- Gold subset: a set of clips labeled by your own experts, with per-layer agreement reported for contact, grasp and active object.
- Side and track integrity: no left/right swaps or identity switches inside a track across occlusions.
- Contact transitions: contact onset and release frames fall within an agreed tolerance of your gold labels; flicker (on-off-on within a few frames) is flagged.
- Mask quality: active-object masks follow thin tools and exclude the hand; hand masks exclude sleeves and gloves if your spec says so.
- Vocabulary drift: object and tool classes match the versioned vocabulary file; free-text classes are rejected.
- Occlusion honesty: occluded keypoints and unseen grasps carry an "occluded" or "unknown" value rather than a guess.
- Annotator provenance: who labeled, with which tool, whether model pre-labels were used and corrected.
The annotation quality audit guide explains sampling sizes, and gold questions and honeypots covers how suppliers should embed checks during labeling. Model-assisted pre-labeling is common for hand masks, so require the disclosure fields described in human annotation provenance.
Where human video stops and robot data starts
Human hand-object video teaches affordances, object roles and task structure, but it does not carry robot joint states or actions. Robot manipulation corpora such as Open X-Embodiment pool demonstrations from robots into a common format for cross-embodiment training [5], and those episodes include the action space human video lacks. Most teams use human video for pre-training or affordance priors and robot data for policy learning.
That split should shape your label budget. If the model will only learn where and how to grasp, contact, active object and grasp type carry most of the value; if it will learn task sequencing, state change and tool relations matter more. The human demonstration video for robot learning guide and the robotics and embodied AI page cover how teams combine the two.
Rights and consent issues specific to hand footage
Hand-object footage still records people, so rights review applies even when no face is visible. Workplace video can capture faces in reflections, coworkers, voices, badges and screens, and tattoos or jewelry on hands can identify a worker. Check the rights layers in a video clip and the notice and consent questions in recording workers on video before you specify capture.
Ask how faces, voices and on-screen text were handled, which method was used, and whether any check was run on a sample. If the footage comes from a clinical or lab setting, health-record rules may apply to anything visible in the frame.
How SourceX fits a hand-object video request
SourceX sources operational datasets, including new recordings of hands-on work, from US companies on request; it does not hold video in stock, and a request does not guarantee a match. SourceX does not source generic CCTV or scraped web video. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and every release is approved by the supplying company. You can describe the footage and label layers you need using the specification above.
Request hand-object interaction video data
Describe the tasks, viewpoint and label layers you need, and SourceX will look for US businesses that hold the data you describe, then run the process from assessment through a license defining records, uses, term and delivery. Nothing is contracted until a supplier agrees. Start a buyer request at sourcex.si/buyers.
Sources
- arXiv (Darkhalil et al.), "EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations" (2022). https://arxiv.org/pdf/2209.13064
- arXiv (Grauman et al.), "Ego4D: Around the World in 3,000 Hours of Egocentric Video" (2021). https://ar5iv.labs.arxiv.org/html/2110.07058
- arXiv / CVPR 2022 (Sener et al.), "Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities" (2022). https://arxiv.org/pdf/2203.14712
- arXiv (Aganian et al.), "ATTACH Dataset: Annotated Two-Handed Assembly Actions for Human Action Understanding" (2023). https://arxiv.org/pdf/2304.08210
- arXiv (Open X-Embodiment Collaboration), "Open X-Embodiment: Robotic Learning Datasets and RT-X Models" (2023). https://arxiv.org/pdf/2310.08864
- CVAT.ai, "COCO (CVAT format documentation)". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
- arXiv / NeurIPS 2021 (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- IEEE Transactions on Human-Machine Systems (Feix et al.), "The GRASP Taxonomy of Human Grasp Types" (2016). https://www.eng.yale.edu/grablab/pubs/Feix_THMS2016.pdf
- CVPR 2020 (Shan et al.), "Understanding Human Hands in Contact at Internet Scale" (2020). https://openaccess.thecvf.com/content_CVPR_2020/papers/Shan_Understanding_Human_Hands_in_Contact_at_Internet_Scale_CVPR_2020_paper.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.