Video data
Pairing Procedure Video with Written SOPs for Step Grounding
Quick answer
A video-to-SOP alignment dataset pairs each recording of a real procedure with the exact written version of the standard operating procedure in force at the time, then maps every SOP step ID to start and end timestamps with a match type: performed, skipped, reordered, or extra. Specify the document version, the boundary tolerance used for agreement, and how deviations are labeled. Without those, a step-grounding or procedural-assistant model learns from silent mismatches between text and action.
By SourceX Editorial · Updated
This page covers the specification problem of aligning a document to video. For buying step-annotated video without a governing document, see procedural and instructional video datasets; for the broader category landscape, start at the video data hub.
Why document-to-video alignment is a different dataset than step-annotated video
Alignment to a written SOP adds a second source of truth, so every label is a relation between a text step and a time span rather than a free-standing action tag. In research corpora, the text side is usually a recipe or a diagram set: one Microsoft Research corpus contains roughly 150K recipe step-to-video alignments across 4,262 dishes [1], and IAW pairs nearly 8,300 IKEA instruction diagrams from 420 furniture types with 1,005 assembly videos [2]. Those sets show the task is well established, but their instructions are public, stable and consumer-grade.
Enterprise SOPs differ in three ways that change the spec. They are versioned and controlled documents, they contain conditional branches ("if torque reading exceeds limit, go to 7.3"), and workers routinely deviate from them for legitimate reasons. Regulated environments formalize this: under 21 CFR 211.100, drug manufacturers must follow written production procedures and record and justify any deviation [4]. A dataset that ignores the deviation record throws away the most informative part of the pairing.
Temporal labels alone are covered on the temporal action segmentation labels page. Here the question is how the text and the timeline reference each other.
The alignment record: fields a step-grounding team should require
The core unit is one alignment row per (video, SOP step, occurrence), keyed to an immutable document version. Without the version key, a revised SOP silently invalidates every timestamp mapped to the old step numbering.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"alignment_id": "aln_000417",
"video_id": "vid_line3_2025-11-04_cam2",
"sop_id": "SOP-ASM-114",
"sop_version": "rev F",
"sop_effective_date": "2025-09-15",
"step_id": "6.2",
"step_text_hash": "sha256:9f1c...",
"occurrence": 1,
"match_type": "reordered",
"start_s": 412.6,
"end_s": 455.1,
"observed_order_index": 5,
"expected_order_index": 6,
"deviation_ref": "DEV-2025-0331",
"annotator_ids": ["a07", "a12"],
"boundary_tolerance_s": 1.0,
"redacted_fields": ["step_6.2_torque_value"],
"view": "egocentric"
}
A few fields carry most of the value:
- sop_version and sop_effective_date. Tie each recording to the revision active on the capture date, not the current one.
- step_text_hash. Detects when step wording changed without a renumbering, a common source of misaligned training pairs.
- occurrence. Rework and retries mean a step can be performed twice; a single start/end pair cannot represent that.
- observed_order_index versus expected_order_index. Lets you compute order violations without re-deriving them from timestamps.
- deviation_ref. Links to the supplier's deviation or nonconformance record when one exists, which is the ground truth for why a step was skipped.
Request the SOP itself as structured text (step IDs, step text, branches, cautions, referenced tools) rather than a scanned PDF. If the source is a PDF, ask how it was parsed and whether step numbering was verified by hand. The dataset delivery formats guide covers manifest and Parquet conventions for shipping the alignment table next to the media.
Labeling skipped, reordered and extra steps as signal
Deviations are the training signal for verification models, so they need first-class labels rather than being cleaned out. A procedural assistant that only ever sees perfect executions cannot learn to say "you skipped the leak test."
Use a closed match-type vocabulary and define each value in the annotation guide:
| match_type | Definition | Has timestamps | Typical cause |
|---|---|---|---|
| performed | Step executed in expected order | Yes | Normal execution |
| reordered | Step executed, but out of expected sequence | Yes | Parallelized work, operator habit |
| skipped | Step in SOP with no corresponding action in video | No | Omission, not applicable, done off-camera |
| extra | Action in video with no SOP step | Yes | Rework, troubleshooting, undocumented practice |
| partial | Step started but not completed | Yes | Interruption, abort |
| not_observable | Step may have occurred but view does not show it | No | Occlusion, camera off |
| not_applicable | Step on a conditional branch that was not taken | No | Branch condition not met |
Illustrative example: invented to show structure; it does not describe an available dataset.
The split between skipped and not_observable is the one most corpora get wrong. A single fixed camera misses many steps; an egocentric view catches hand actions but loses context. Research on aligning sparse instructions to first-person video uses egocentric cues such as hand and gaze information to localize fine-grained actions [3]. Ask suppliers which view each step label was made from, and whether any label was inferred from a paper checklist rather than seen on video.
Conditional branches need their own handling. Record the branch taken (for example, "7.3 via 6.4 fail") so a skipped step on the untaken branch is labeled not_applicable rather than skipped.
Measuring boundary agreement with a stated tolerance
Report inter-annotator agreement on step boundaries as a function of an explicit time tolerance, because "agreement" without a tolerance is not comparable across datasets. Ask for double annotation on a sample of videos and the tolerance windows used, such as 0.5, 1 and 2 seconds.
Two metrics are worth requesting separately:
- Match-type agreement. A categorical label per (video, step), so use a chance-corrected statistic. Cohen's kappa suits two annotators; Krippendorff's alpha handles more annotators and missing labels [5].
- Boundary agreement. For steps both annotators marked as present, report the share of start and end boundaries within tolerance, plus temporal IoU of the segments.
Boundary disagreement concentrates at step transitions where the SOP is ambiguous ("prepare the fixture") rather than at crisp physical events ("press cycle start"). Ask for per-step agreement so you can drop or down-weight steps whose boundaries are inherently fuzzy. For a broader framework, ISO/IEC 5259-3 sets requirements for a data quality management process for ML data without fixing specific metrics [6].
Confidential procedures, redaction and worker privacy
Real SOPs often encode proprietary process parameters, so expect supplier-approved redaction of values such as torque specs, temperatures, recipe quantities and part numbers. Redaction must be recorded per field so your model knows a value is masked, not absent.
Common redaction patterns and their effect on training:
- Text-side masking. Replace parameter values in step text with typed placeholders like
<TORQUE_NM>. Grounding still works; numeric verification does not. - Video-side masking. Blur displays, HMI screens or labels that show the same values. Check that blurs do not cover the hands, or step boundaries become unlabelable.
- Step withholding. Remove whole steps. The alignment table must still list them as withheld, or skipped-step statistics become biased.
The people in the video are a separate rights layer. Faces and hands are visible in most procedure footage, and some state laws treat a record of hand or face geometry as a biometric identifier requiring notice and consent before commercial capture [8]. The video clip rights layers and recording workers on video pages cover consent, audio and notice questions in detail. On the document side, data statements give a structured way to record who produced the SOP text and the annotations [7].
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Where SOP-paired procedure video comes from
The most realistic supply is organizations that already control both artifacts: a document control system holding versioned work instructions, and video captured for training, quality or new recordings of hands-on work. Assembly, maintenance, lab and warehouse operations are the usual candidates; see the manufacturing assembly video and assembly demonstrations with work instructions and CAD pages for adjacent specs.
SourceX sources operational datasets from US companies on request, including documents, engineering records and new recordings of hands-on work; nothing is held in stock and a request does not guarantee a match. Buyers describe the data, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded, and delivery happens through private, access-controlled workflows after an executed agreement. You can describe the procedure video and SOP pairing you need. For the SOP text on its own, see SOP and knowledge base datasets and licensing SOPs and playbooks; for the footage side, see egocentric video of skilled manual work.
Request checklist for an SOP-to-video alignment dataset
Use this checklist to turn a research need into a supplier-ready request.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Procedure domain and step count range (for example, 15 to 60 steps per SOP)
- SOP delivered as structured text with step IDs, branches and cautions
- Revision history with effective dates for every SOP version referenced
- One alignment row per (video, step, occurrence), with match_type from a closed vocabulary
- Explicit not_observable and not_applicable labels separate from skipped
- Links to deviation or nonconformance records where they exist
- Camera views per video (egocentric, fixed, multi-view) and sync method
- Double annotation on a stated sample, with kappa or alpha and boundary agreement at named tolerances
- Per-field redaction log for proprietary parameters and on-screen values
- Worker consent and notice basis, and de-identification method for faces and voices
- Held-out split by site or SOP, not by clip, to prevent leakage of the same procedure
- License terms covering allowed uses, term and delivery for both the video and the document text
The leakage item matters for evaluation. Splitting by clip lets the same SOP and operator appear in train and test, inflating step-localization scores. Split by SOP ID or site, and keep a slice of deviation-heavy videos in the test set.
For related evaluation and retrieval uses, see the AI data hub and the RAG content licensing guide.
Sourcing an SOP-paired procedure video dataset
SourceX serves AI teams wherever they are based and runs each request through Find, Assess, Agree, Transact and Manage, with data and licensing permissions assessed before any license on pricing and allowed uses is agreed. Nothing is contracted until a supplier agrees. To describe the procedures, views and alignment fields you need, start a buyer request.
Sources
- arXiv (Microsoft Research), "A Recipe for Creating Multimodal Aligned Datasets for Sequential Tasks" (2020). https://arxiv.org/pdf/2005.09606
- arXiv (CVPR 2023), "Aligning Step-by-Step Instructional Diagrams to Video Demonstrations" (2023). https://arxiv.org/html/2303.13800v4
- arXiv (ar5iv), "Learning to Localize and Align Fine-Grained Actions to Sparse Instructions" (2018). https://ar5iv.arxiv.org/html/1809.08381
- Legal Information Institute, Cornell Law School, "21 CFR 211.100 - Written procedures; deviations". https://www.law.cornell.edu/cfr/text/21/211.100
- arXiv, "Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation" (2026). https://arxiv.org/pdf/2603.06865
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-3:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 3: Data quality management requirements and guidelines" (2024). https://www.iso.org/standard/81092.html
- University of Washington Tech Policy Lab, "Data Statements" (2021). https://techpolicylab.uw.edu/data-statements/
- Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier" (2026). https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.