Skip to content

Video data

Procedural and Instructional Video Datasets with Step-Level Annotations

Quick answer

A procedural video dataset pairs recordings of a task being performed with a task label and an ordered list of steps, each with start and end timestamps and step text. For commercial training, the most useful supply is unedited footage of real work, recorded with worker consent and paired with the written procedure, including the deviations, retries and corrections that public benchmarks often trim away. Specify viewpoint, annotation layers and rights up front, then license from the organization that owns the footage.

By SourceX Editorial · Updated

How licensed procedural video differs from public instructional benchmarks

Public procedural benchmarks are excellent for research comparison, but most were not designed as commercial training supply. Assembly101 records toy-vehicle assembly and disassembly from 8 static and 4 egocentric cameras, with over 100K coarse and 1M fine-grained action segments plus mistake labels [1]. COM Kitchens uses unedited overhead smartphone video of people following recipes, linked to instructions through visual action graphs [3]. Newer benchmarks such as ProcObject-10K test object-centric question answering and mistake recognition over instructional video [2].

Several widely used instructional sets, including COIN, CrossTask and HowTo100M, are assembled from online how-to videos [11]. Before using any of them in a commercial model, read the dataset license and the terms attached to the underlying videos, because the annotation license and the footage rights are often separate questions. Ego4D, with more than 3,000 hours of egocentric video [4], carries its own license agreement [10]; check it rather than assuming research access extends to product training.

The practical gap is domain and realism. Web how-to video is usually staged, edited and narrated for a viewer, so steps appear in tidy order with cuts over the boring parts. Operational video from a repair bench, a lab, a kitchen line or a service call shows idle time, tool searches, re-dos and steps done out of order, which is exactly what step-recognition and task-guidance models meet in deployment.

The annotation layers to specify in a step-annotated video request

A useful step-annotated dataset has at least five layers, and each one should be named in your request. Without them, vendors default to whatever their labeling tool exports, which is usually coarse clip-level tags.

  • Task label: one task per recording (for example, "replace HVAC blower motor"), drawn from a controlled vocabulary you supply.
  • Ordered step segments: step ID, start and end time in milliseconds against the source video's own timeline, and the step's position in the canonical procedure.
  • Step text: a short imperative description, plus a pointer to the matching line of the written procedure when one exists.
  • Required versus optional flags: which steps the procedure mandates and which are conditional (for example, "bleed line if pressure is above spec").
  • Deviation and error events: skipped steps, out-of-order steps, wrong-tool or wrong-part errors, and the corrective action, each with timestamps.

Fine-grained action labels, as in Assembly101 [1], are a separate and more expensive layer; ask for them only if your model needs sub-step action recognition. For segment-level label design and boundary rules, see the guide to temporal action segmentation and localization labels.

Keep deviations from the SOP instead of cleaning them out

Real operations rarely follow the written order, and a dataset that hides this teaches a model the wrong distribution. Regulated industries already treat deviations as records: US drug manufacturing rules require that written production procedures be followed and that any deviation be recorded and justified [5]. Ask suppliers to annotate deviations as events rather than discard or re-shoot those takes.

Deviations are also your best evaluation material. A held-out split weighted toward out-of-order and corrected runs tests whether a task-guidance assistant can recover gracefully. For a deeper look at sourcing error examples, see mistake and deviation examples in procedural video.

Choosing a viewpoint: fixed, over-the-shoulder or egocentric

Viewpoint determines what a model can learn, so choose it from the deployment camera, not from what is easiest to record. Fixed third-person cameras suit workstation monitoring and quality checks; over-the-shoulder views capture hands and the work surface with context; head-mounted capture matches assistants that run on smart glasses.

Multi-view capture, as in Assembly101's synchronized static and egocentric rigs [1], helps when you want to train view-invariant step recognition, but it multiplies storage and annotation cost. Ask for camera intrinsics, frame rate, resolution and a shared clock across views if you take more than one. For head-mounted manual work specifically, SourceX maintains a dedicated page on egocentric video datasets of skilled manual work, and the egocentric video glossary entry defines the term.

Pair each recording with its written procedure

The highest-value procedural video comes with the document the worker was meant to follow. A recording linked to a versioned SOP, work instruction or checklist lets you train step grounding, generate step text from the source document, and measure how often practice diverges from the written steps. COM Kitchens shows the research value of linking instructions to video at the action level [3].

Ask for the SOP version ID in each recording's metadata, since procedures change and a mismatch silently corrupts alignment labels. The video-to-SOP step alignment guide covers alignment formats, and SourceX also lists SOPs and playbooks as a document category.

Rights checks before you license procedural footage

Procedural footage of real work carries several rights layers, and each must clear before the data trains a commercial model. The footage owner (usually the employer) must hold the recording and have authority to license it. Workers in frame should receive notice and give consent that covers AI training, and face or hand geometry derived from video can trigger biometric statutes such as Illinois BIPA [6] and Texas Business and Commerce Code Section 503.001 [7].

Look also for proprietary equipment, customer documents, screens and labels in frame, and for audio that captures conversation. The breakdown of rights layers in a video clip and the guide to recording employees on video for AI datasets go deeper. Request a datasheet in the style of Data Cards, covering collection method, annotation process and intended use [8].

Request template for a procedural video dataset

The fastest way to get comparable answers from suppliers is to send the same structured request to each. Adapt the record below to your task vocabulary and deliver annotations as JSON Lines, one recording per line, UTF-8 encoded [9]. The record below is wrapped for readability; in the file it occupies a single line.

Illustrative example: invented to show structure; it does not describe an available dataset.

{"recording_id": "rec_000184", "task_label": "replace_blower_motor", "sop_id": "WI-HV-022", "sop_version": "4.1",
 "viewpoint": "over_the_shoulder", "fps": 30, "resolution": "1920x1080", "duration_ms": 1265000,
 "steps": [
  {"step_id": "s1", "canonical_index": 1, "start_ms": 4200, "end_ms": 61800, "text": "Disconnect power at service switch", "required": true},
  {"step_id": "s3", "canonical_index": 3, "start_ms": 61800, "end_ms": 190400, "text": "Remove blower access panel", "required": true},
  {"step_id": "s2", "canonical_index": 2, "start_ms": 190400, "end_ms": 214000, "text": "Verify zero voltage with meter", "required": true}
 ],
 "deviations": [{"type": "out_of_order", "step_id": "s2", "start_ms": 190400, "note": "Voltage check done after panel removal"}],
 "consent": {"worker_notice": true, "ai_training_scope": true}, "faces_blurred": false, "audio_present": true}

Pair the template with a short checklist for each supplier:

QuestionWhy it matters
Is the footage unedited, with original timestamps?Edited video hides idle time and breaks boundary labels
Which written procedure, and which version, applies to each recording?Enables step grounding and deviation labels
How were step boundaries defined and adjudicated?Boundary rules drive inter-annotator agreement
Are deviations, retries and corrections labeled or removed?Removed errors bias models toward ideal order
What consent did workers give, and does it cover AI training?Determines whether the footage can be licensed
Are faces, screens, documents or audio in frame?Drives redaction scope and biometric exposure

Where procedural video fits among other video data

Step-annotated procedural video sits between general action recognition and video-language instruction data. If your target is assembly-line action classes, the manufacturing assembly video datasets page is closer; if you need question-answer pairs over video for a video LLM, see video question-answer and instruction data. Existing internal training libraries are covered in licensing corporate training and how-to video libraries, and the video data hub maps the full cluster.

SourceX sources operational datasets from US companies, including new recordings of hands-on work, and manages the licensing process; the buyer intake is where you describe the task, viewpoint and annotation layers. Supply is sourced on request rather than held in stock, so a request does not guarantee a match. See also the video recordings category and the AI data hub.

Request procedural video with step annotations

SourceX looks for US businesses that hold the procedural video you describe, reviews ownership and consents, and delivers only under a license that defines records, uses, term and delivery, with every release approved by the supplying company. Personal details are removed or replaced before delivery, with the method recorded and a sample checked. Describe your task vocabulary, viewpoint and annotation layers at sourcex.si/buyers.

Sources

  1. CVF Open Access (CVPR 2022), "Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities" (2022). https://openaccess.thecvf.com/content/CVPR2022/html/Sener_Assembly101_A_Large-Scale_Multi-View_Video_Dataset_for_Understanding_Procedural_Activities_CVPR_2022_paper
  2. arXiv, "ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos" (2025). https://arxiv.org/pdf/2512.03479
  3. ML Anthology, "COM Kitchens (ECCV 2024)" (2024). https://mlanthology.org/eccv/2024/hashimoto2024eccv-com
  4. arXiv, "Ego4D: Around the World in 3,000 Hours of Egocentric Video" (2021). https://arxiv.org/abs/2110.07058v1
  5. Legal Information Institute, Cornell Law School, "21 CFR 211.100 - Written procedures; deviations". https://www.law.cornell.edu/cfr/text/21/211.100
  6. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  7. Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  8. Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  9. jsonlines.org, "JSON Lines". https://jsonlines.org/
  10. Ego4D Consortium, "Ego4D License Agreement" (2026). https://ego4d-data.org/pdfs/Ego4D-Licenses-Draft.pdf
  11. Tang et al. (Tsinghua University / Meitu Inc.), "COIN Dataset" (2019). https://github.com/coin-dataset/annotations/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data