Video data
Temporal Action Segmentation and Localization Labels: Specifying Annotated Video
Quick answer
A temporal action segmentation dataset gives every frame of an untrimmed video an action label, including background, so a model learns where actions start, stop and follow each other. Localization labels mark only sparse start-end segments of interest, and recognition labels name pre-trimmed clips. Before buying, fix the label type to your evaluation metric, then specify a hierarchical taxonomy, written boundary rules with a tolerance, a background class, and agreement measured separately for boundaries and class names.
By SourceX Editorial · Updated
Segmentation, localization or recognition: which label type to buy
Buy the label type your evaluation metric consumes, because each one is scored differently and costs a different amount of annotator time. Dense segmentation is scored with frame-wise accuracy, segmental edit score and segmental F1 at 10%, 25% and 50% overlap, the convention used on academic benchmarks such as Breakfast, 50Salads and GTEA. Localization is scored with mean average precision over temporal IoU (tIoU) thresholds; ActivityNet-style evaluation averages mAP over strict thresholds from 0.5 to 0.95, while THUMOS14-style evaluation uses lower ones. Recognition needs only a class per trimmed clip, which is the cheapest to annotate but teaches nothing about boundaries.
| Label type | Unit annotated | Typical scoring | Relative annotation effort | Use it for |
|---|---|---|---|---|
| Temporal action segmentation | Every frame, background included | Frame accuracy, edit score, F1@{10,25,50} | Highest: annotator must watch the full video and close every gap | Procedural monitoring, step tracking, assembly verification |
| Temporal action localization | Sparse start-end segments per class | mAP averaged over tIoU thresholds | Medium: unlabeled spans are implicitly negative | Event retrieval, highlight detection, video-LLM grounding |
| Action recognition | One label per trimmed clip | Top-1 / top-5 accuracy | Lowest, but trimming already decided the boundaries | Classifier pretraining, taxonomy validation |
A common mistake is buying localization labels and then training a segmentation model. Unlabeled spans in a localization file mean "not one of our classes," not "idle," so converting them to background silently injects noise. If you need both, ask for dense segmentation; sparse segments can always be derived from it, but not the reverse.
Designing the action taxonomy for untrimmed video
A usable action recognition label taxonomy is hierarchical, mutually exclusive at each level, and written as verb-object pairs at the finest level. Assembly101 shows the pattern at scale: more than 100K coarse segments and 1M fine-grained segments across synchronized static and egocentric views [1]. A practical three-level hierarchy is activity (for example, "replace filter cartridge"), action or step ("remove housing"), and verb-object atom ("unscrew housing-cap"). Annotate the finest level and roll up, since coarse labels cannot be split later without re-watching the footage.
Long-tail classes need an explicit policy. Decide up front whether rare verb-object pairs get their own class, merge into a parent ("manipulate housing"), or fall into an "other-action" bucket, and require the supplier to report per-class segment counts and total seconds. Frame accuracy is dominated by long classes, so a dataset can score well while a short, rare and safety-relevant action is nearly absent.
For procedure footage, tie the taxonomy to the written procedure; our guides on procedural and instructional video datasets and pairing procedure video with written SOPs cover step-level alignment. Factory-floor taxonomies are covered in manufacturing assembly video datasets.
Boundary rules and tolerance: the part guidelines usually skip
Action boundary annotation needs written start and end rules per class, because annotators disagree more on when an action starts than on what it is. Define the start as a visible event ("hand contacts the tool") rather than intent ("worker decides to reach"), and define the end as the first frame of the release or the next action. Segmental edit and F1 metrics deliberately tolerate small timing shifts while penalizing over-segmentation, which tells you how much boundary precision a model actually needs.
Set a boundary tolerance in seconds or frames per class (for example, plus or minus 0.5 s for fast hand actions, 2 s for coarse steps) and measure boundary agreement separately from label agreement. Two annotators can agree perfectly on the class sequence while their boundaries drift by a second; a single blended agreement number hides that. Over-segmentation, where one action is split into flickering fragments, is a well-documented failure mode in segmentation models, so guidelines should also set a minimum segment duration and say how to merge micro-pauses.
Specify the frame rate the labels refer to. Labels written at 15 fps and delivered against 30 fps source video produce off-by-one-frame errors at every boundary unless timestamps are in seconds or the frame rate is stated in the manifest.
Background, idle and transition handling
Every dense segmentation label set needs an explicit background class with a definition, not an unlabeled gap. Split it if the distinction matters to the model: "idle" (no task activity), "transition" (moving between steps) and "occluded or off-camera" (the action may be happening but cannot be seen). Merging these into one background label teaches the model that an occluded weld and a coffee break look the same.
Ask for the background share per video. A dataset where 60% of frames are background behaves very differently from one where 10% are, and the supplier should report it rather than leave you to discover it after training. Overlapping actions (two hands doing different things) need a rule too: either a primary-action convention or separate label tracks per hand or actor.
Layered QA and agreement you can audit
Ask for layered QA: annotator qualification, hidden gold checks, and multi-annotator consensus on a sample, with inconclusive items flagged rather than forced into a class. The EPIC-KITCHENS VISOR effort, which produced pixel-level object segmentations rather than temporal labels, describes a comparable multi-stage annotation and checking pipeline [2]. Merging several annotators' segment lists into one label track is harder than majority vote on a single class, and aggregation methods for complex, structured annotations exist for this reason [3]. For class agreement on segments, Krippendorff's alpha handles more than two annotators and missing ratings [4]; for boundaries, report the distribution of start and end offsets between annotators in seconds. Our guide to inter-annotator agreement metrics explains the trade-offs, and annotation adjudication covers how disagreements become final labels.
Do not assume benchmark-grade data is clean. Audits of widely used test sets estimated an average label error rate of at least 3.3% [5], and temporal labels have two error surfaces (class and boundary) instead of one. Reserve a held-out slice for your own re-annotation before accepting a full delivery.
Require documentation of how labels were produced, including guideline version, annotator pool, tool, and whether model pre-labels were used, in a Data Card style summary [6]. Provenance for human annotation is covered in human annotation provenance.
A request specification you can send to suppliers
A complete request names the label type, taxonomy, boundary rules, background policy, QA layers and acceptance tests in one document. If you are licensing existing labels rather than commissioning new ones, the same fields apply; see licensing expert annotations and labels. The template below is one way to structure it; adapt the fields to your model.
Illustrative example: invented to show structure; it does not describe an available dataset.
label_spec:
label_type: dense_temporal_segmentation # or: localization | recognition
time_basis: seconds # never bare frame indices
source_fps_declared: true
taxonomy:
levels: [activity, step, verb_object]
annotate_at: verb_object # roll up to coarser levels
long_tail_policy: "classes under 50 segments merge to parent; report counts"
boundaries:
start_rule: "first frame of visible hand-object contact"
end_rule: "first frame of release or next action"
tolerance_s: { verb_object: 0.5, step: 2.0 }
min_segment_s: 0.3
background:
classes: [idle, transition, occluded_or_offscreen]
overlaps: per_hand_tracks # or: primary_action_only
qa:
qualification_test: required
gold_check_rate: "5% of segments, hidden"
multi_annotator_sample: "10% of videos, 3 annotators"
class_agreement: krippendorff_alpha
boundary_agreement: "median and p90 start/end offset in seconds"
inconclusive_flag: allowed
acceptance:
buyer_reannotation_slice: "2% of videos"
report: [per_class_counts, per_class_seconds, background_share, guideline_version]
Annotation file encoding (JSON, CSV per frame, or tool-native exports) is a delivery question; keep this spec focused on what the labels mean.
Rights questions that come with labeled workplace video
Labels do not change the rights in the underlying footage, so clear the video before you pay for annotation. Footage of employees, customers or patients carries consent and biometric questions; see recording employees on video and rights layers in a video clip. Start from the video data hub for collection sources and checks, and see data annotation in the glossary for terminology.
When SourceX handles a request, each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe the labeled video you need on the buyers page.
Sourcing temporal action segmentation data
SourceX sources operational datasets, including new recordings of hands-on work, from US companies on request; it holds no stock, and a request does not guarantee a match. Each release is approved by the supplying company, and nothing is contracted until a supplier agrees. Describe your temporal labeling requirements on the SourceX buyers page.
Frequently asked questions
Can I convert localization labels into segmentation labels?
Only partially. Unlabeled spans in a localization file are undefined rather than background, so a conversion produces a background class that mixes idle time with unannotated actions. Ask for dense labels if segmentation is the goal.
What frame rate should action labels use?
Store boundaries in seconds and declare the source frame rate. That lets you resample video for feature extraction without shifting every boundary.
Is one agreement number enough to accept a dataset?
No. Report class agreement (for example, Krippendorff's alpha on segments) and boundary agreement (offsets in seconds) separately, plus per-class counts so rare actions are visible.
Sources
- arXiv (Sener et al., CVPR 2022), "Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities" (2022). https://arxiv.org/pdf/2203.14712
- arXiv, "EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations" (2022). https://arxiv.org/pdf/2209.13064
- arXiv, "A General Model for Aggregating Annotations Across Simple, Complex, and Multi-Object Annotation Tasks" (2023). https://arxiv.org/pdf/2312.13437
- University of Pennsylvania, Annenberg School for Communication, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
- arXiv (Northcutt, Athalye, Mueller; NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- arXiv (Pushkarna, Zaldivar, Kjartansson; FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.