Skip to content

Video data

Hand-Object Interaction Annotation in Video: Grasps, Contacts and Tool Use

Quick answer

A useful hand-object interaction dataset labels more than hands. For manipulation learning you need, per hand and per frame or segment: hand side and location, contact state, the active object (ideally as a mask), grasp type from a fixed taxonomy, the object's state change, and the tool-to-target relation when a tool acts on something else. Specify each layer, its sampling rate and its QA rule before collection, because these labels are hard to retrofit onto footage later.

By SourceX Editorial · Updated

This page is a label specification. For the footage itself and how it is captured, see the video data hub and the egocentric video of skilled manual work page; for verb-level time segments, see temporal action segmentation labels.

Which label layers a manipulation team actually needs

Most manipulation-oriented teams need six layers, and each one answers a different modeling question. Public benchmarks show that the layers stack: EPIC-KITCHENS VISOR, for example, pairs pixel-level masks for hands and active objects in egocentric kitchen video with hand-object relations [1]. A buyer spec should say which layers are mandatory and which are optional.

  • Hand detection and side. Box or mask per hand, tagged left or right, with a persistent track ID across the clip. Without the side tag you cannot model bimanual role split.
  • Contact state. At minimum no-contact, self-contact, other-person contact, portable-object contact and fixed-object contact [9]. Fixed versus portable matters for robotics, because pushing a drawer and lifting a cup need different policies.
  • Active object. The object the hand is manipulating, as a box, polygon or mask, linked to the hand by ID. VISOR's choice of masks over boxes reflects that thin tools and partially occluded objects are poorly described by rectangles [1].
  • Grasp type. A class from a published grasp taxonomy, such as the GRASP taxonomy of static, one-handed grasps organized by opposition type, power versus precision and thumb position [8], or a coarser merged version of it.
  • Object state change. Pre-condition, post-condition and the frame where the change occurs (open to closed, whole to cut, loose to fastened).
  • Tool-to-target relation. For tool use, a triple of hand, tool and target object with the effect on the target, for example screwdriver acting on screw, screw becoming seated.

Hand keypoints (2D joints) are a seventh layer that some teams add. Once you add depth, IMU or 3D hand pose and mesh, the dataset becomes a multimodal capture problem with calibration and synchronization requirements, which is a different specification from video labels alone.

How to pick a grasp taxonomy annotators can apply consistently

Pick the coarsest taxonomy your model can use, because inter-annotator agreement drops quickly as classes multiply. A full fine-grained grasp taxonomy is precise but hard to apply from a single camera angle where fingers are occluded. Many teams label a merged set of general grasp classes, or a three-way power, precision and intermediate split, and reserve fine classes for a calibrated subset.

Two rules prevent the most common confusion. First, label grasp only while contact state is portable-object contact and the grasp is stable, since grasp taxonomies describe static, stable grasps and transitional finger motion does not fit them. Second, add an explicit "non-prehensile" value for pushing, pressing and sliding, so annotators are not forced to pick the nearest grasp class for a hand that is not grasping.

Industrial work needs a domain vocabulary on top of the grasp classes. A general taxonomy will not tell you whether a hand holds a torque wrench, a crimper or a pipette; those come from a tool and part list agreed with domain reviewers before labeling starts. The manufacturing assembly video guide covers how assembly vocabularies are usually built.

How viewpoint changes occlusion, label cost and what the model learns

Egocentric and third-person video produce different labels for the same task, so specify viewpoint before you specify labels. Head-mounted footage, as in Ego4D's more than 3,000 hours of first-person video [2], keeps hands near the image center and large, which makes contact and active-object labels easier, but the wrist and the far hand often leave the frame. Fixed third-person cameras see both hands and the body, which helps bimanual role labels, but small tools shrink to a few pixels.

Multi-view capture resolves many occlusion disputes. Assembly101 recorded eight static and four egocentric views simultaneously of people taking apart and rebuilding toy vehicles [3], which lets an annotator check a hidden grasp in another view. If you buy multi-view footage, require frame-level synchronization metadata, or cross-view label propagation will drift.

Bimanual tasks need per-hand labels, not per-person labels. ATTACH annotated assembly actions by hand, including two-handed and simultaneous actions [4], because a single "person is screwing" label hides the stabilizing hand that a robot also has to reproduce. If your target platform is dual-arm, make per-hand labeling a hard requirement.

An illustrative label specification and record

The practical deliverable is a label specification that names every layer, its geometry, its sampling rate and its acceptance rule. Writing it down also forces agreement on what "contact" and "active object" mean before anyone draws a mask.

Illustrative example: invented to show structure; it does not describe an available dataset.

LayerGeometry and valuesSamplingAcceptance check
Hand trackMask or box, side L/R, track_idEvery annotated frameSide swaps per track = 0 in audited clips
Contact statenone, self, person, portable_obj, fixed_objEvery annotated frameAgreement with second annotator on a gold subset
Active objectMask (polygon or RLE), object_class, object_idFrames with contactMask IoU against expert gold above an agreed threshold
Grasp typeMerged general grasp classes plus non_prehensileStable contact segments onlyConfusion matrix reviewed per class
Hand keypoints21 2D joints, visibility flag per jointKeyframesOccluded joints flagged, not guessed
State changepre_state, post_state, change_framePer interaction segmentchange_frame within a tolerance window
Tool relationtool_id, target_id, effect verbPer tool-use segmentTarget linked to an existing object_id

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "clip_id": "clip_000412",
  "frame": 1834,
  "view": "egocentric_head",
  "hands": [
    {
      "track_id": "h1",
      "side": "right",
      "bbox": [612, 388, 142, 120],
      "contact_state": "portable_obj",
      "active_object_id": "o7",
      "grasp_type": "power_medium_wrap",
      "keypoints_visible": 17
    },
    {
      "track_id": "h2",
      "side": "left",
      "bbox": [301, 402, 130, 118],
      "contact_state": "fixed_obj",
      "active_object_id": "o2",
      "grasp_type": "non_prehensile"
    }
  ],
  "objects": [
    {"object_id": "o7", "class": "screwdriver", "segmentation": "rle:...", "role": "tool"},
    {"object_id": "o2", "class": "panel_bracket", "segmentation": "rle:...", "role": "target"}
  ],
  "interaction": {
    "tool_id": "o7",
    "target_id": "o2",
    "effect": "fasten",
    "pre_state": "screw_loose",
    "post_state": "screw_seated",
    "change_frame": 1902
  }
}

Ask for an export your training stack already reads. COCO-style JSON keeps images, categories and annotations in separate sections and supports polygon or RLE masks and keypoint tasks [6], which covers hands, objects and joints; contact, grasp and relation fields then go in per-annotation attributes or a sidecar file keyed by annotation ID. Whatever the format, require a schema document and a versioned label vocabulary file with the delivery.

How to check hand-object labels before you accept a delivery

Treat hand-object labels as suspect until you have audited a sample, because even widely used benchmarks carry measurable label noise. An audit of 10 popular test sets estimated an average label error rate of at least 3.3% [7], and contact and grasp labels are more subjective than most image classes. VISOR's authors used a multi-stage annotation and checking process to keep masks consistent over time [1], which is the kind of process detail to ask a supplier to describe.

Illustrative example: invented to show structure; it does not describe an available dataset.

Acceptance checklist for a hand-object annotation delivery

  1. Gold subset: a set of clips labeled by your own experts, with per-layer agreement reported for contact, grasp and active object.
  2. Side and track integrity: no left/right swaps or identity switches inside a track across occlusions.
  3. Contact transitions: contact onset and release frames fall within an agreed tolerance of your gold labels; flicker (on-off-on within a few frames) is flagged.
  4. Mask quality: active-object masks follow thin tools and exclude the hand; hand masks exclude sleeves and gloves if your spec says so.
  5. Vocabulary drift: object and tool classes match the versioned vocabulary file; free-text classes are rejected.
  6. Occlusion honesty: occluded keypoints and unseen grasps carry an "occluded" or "unknown" value rather than a guess.
  7. Annotator provenance: who labeled, with which tool, whether model pre-labels were used and corrected.

The annotation quality audit guide explains sampling sizes, and gold questions and honeypots covers how suppliers should embed checks during labeling. Model-assisted pre-labeling is common for hand masks, so require the disclosure fields described in human annotation provenance.

Where human video stops and robot data starts

Human hand-object video teaches affordances, object roles and task structure, but it does not carry robot joint states or actions. Robot manipulation corpora such as Open X-Embodiment pool demonstrations from robots into a common format for cross-embodiment training [5], and those episodes include the action space human video lacks. Most teams use human video for pre-training or affordance priors and robot data for policy learning.

That split should shape your label budget. If the model will only learn where and how to grasp, contact, active object and grasp type carry most of the value; if it will learn task sequencing, state change and tool relations matter more. The human demonstration video for robot learning guide and the robotics and embodied AI page cover how teams combine the two.

Hand-object footage still records people, so rights review applies even when no face is visible. Workplace video can capture faces in reflections, coworkers, voices, badges and screens, and tattoos or jewelry on hands can identify a worker. Check the rights layers in a video clip and the notice and consent questions in recording workers on video before you specify capture.

Ask how faces, voices and on-screen text were handled, which method was used, and whether any check was run on a sample. If the footage comes from a clinical or lab setting, health-record rules may apply to anything visible in the frame.

How SourceX fits a hand-object video request

SourceX sources operational datasets, including new recordings of hands-on work, from US companies on request; it does not hold video in stock, and a request does not guarantee a match. SourceX does not source generic CCTV or scraped web video. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and every release is approved by the supplying company. You can describe the footage and label layers you need using the specification above.

Request hand-object interaction video data

Describe the tasks, viewpoint and label layers you need, and SourceX will look for US businesses that hold the data you describe, then run the process from assessment through a license defining records, uses, term and delivery. Nothing is contracted until a supplier agrees. Start a buyer request at sourcex.si/buyers.

Sources

  1. arXiv (Darkhalil et al.), "EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations" (2022). https://arxiv.org/pdf/2209.13064
  2. arXiv (Grauman et al.), "Ego4D: Around the World in 3,000 Hours of Egocentric Video" (2021). https://ar5iv.labs.arxiv.org/html/2110.07058
  3. arXiv / CVPR 2022 (Sener et al.), "Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities" (2022). https://arxiv.org/pdf/2203.14712
  4. arXiv (Aganian et al.), "ATTACH Dataset: Annotated Two-Handed Assembly Actions for Human Action Understanding" (2023). https://arxiv.org/pdf/2304.08210
  5. arXiv (Open X-Embodiment Collaboration), "Open X-Embodiment: Robotic Learning Datasets and RT-X Models" (2023). https://arxiv.org/pdf/2310.08864
  6. CVAT.ai, "COCO (CVAT format documentation)". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
  7. arXiv / NeurIPS 2021 (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  8. IEEE Transactions on Human-Machine Systems (Feix et al.), "The GRASP Taxonomy of Human Grasp Types" (2016). https://www.eng.yale.edu/grablab/pubs/Feix_THMS2016.pdf
  9. CVPR 2020 (Shan et al.), "Understanding Human Hands in Contact at Internet Scale" (2020). https://openaccess.thecvf.com/content_CVPR_2020/papers/Shan_Understanding_Human_Hands_in_Contact_at_Internet_Scale_CVPR_2020_paper.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data