Video data
Mistake and Deviation Examples in Procedural Video: Sourcing Real Error Data
Quick answer
A usable mistake detection video dataset pairs long recordings of real work with time-stamped error segments typed by kind (omission, wrong order, wrong part or tool, technique error) and the corrections that follow. Public benchmarks such as Assembly101 show the label structure, but most induce or script errors in lab settings. Natural errors are rare, so buyers should plan for targeted sampling and for labeling errors after the fact by joining video to rework and nonconformance records.
By SourceX Editorial · Updated
This page covers the error-specific purchase. For step-annotated procedural video in general, start with procedural and instructional video datasets with step-level annotations; for the wider category, see the video data hub.
Staged versus natural mistakes: check how each benchmark induced its errors
The first question for any error dataset is how the errors got there, because staged and natural mistakes have different visual signatures and different frequencies. A participant told to skip a step tends to skip it cleanly; a tired technician who forgets a washer often hesitates, glances back, or corrects later. Models trained only on scripted errors can learn the cue of the script rather than the error itself.
The public benchmarks differ on exactly this point:
- Assembly101 records people disassembling and reassembling take-apart toy vehicles from 8 static and 4 egocentric views, with over 100K coarse and 1M fine-grained action segments, and includes mistake detection among its benchmarks [1]. Because participants work without step-by-step instructions, its mistakes arise during unscripted assembly, but the domain is a toy, not a production line.
- Other error-focused benchmarks in cooking, guided manipulation and industrial-like assembly mix recipe-following runs with runs where participants were asked to deviate, or capture mistakes under live remote instruction, where an instructor's intervention often marks the error. Read each collection protocol to see what share of errors was induced.
- Class imbalance is the common pattern: even deliberately designed error sets typically contain far fewer incorrect step completions than correct ones, so check per-class counts in each split before relying on a reported score.
Use these to fix your schema and baseline, then check each license before using them for anything beyond research. None of them is a substitute for errors made by trained workers on real equipment.
An error taxonomy that survives contact with real work
Real procedural errors fall into a small set of types, and the taxonomy should be fixed before any footage is labeled. Without it, one annotator's "wrong order" is another's "omission plus late insertion," and inter-annotator agreement collapses on exactly the rare classes you are paying for.
A practical starting taxonomy:
- Omission: a required step never happens (torque check skipped, gasket not installed).
- Wrong order: all steps occur, but a dependency is violated (fastener tightened before alignment).
- Wrong part or tool: the step happens with the wrong component, fastener length, bit or consumable.
- Technique error: right part, right order, wrong execution (cross-threading, uneven pressure, contamination).
- Correction: the worker notices and fixes the error; label its start, end and which error it resolves.
- Near-miss: the worker begins an error and aborts it before it takes effect.
Corrections and near-misses matter most for quality-assist and task-verification models, because they teach when to intervene and when the worker has already caught it. They are also what scripted datasets most often lack.
Errors are rare: size the recording plan from the error rate, not the hours
Natural mistake footage is scarce, so dataset size should be planned backward from the number of error events you need per class. If a skilled line produces a few logged defects per hundred units, thousands of normal cycles are needed to collect a few dozen examples of any one error type, and the long tail of technique errors may never reach a trainable count.
Three practical responses:
- Target sampling windows. Ask suppliers to prioritize shifts, stations, new-hire periods, or product changeovers where their quality records show higher defect rates.
- Keep the normal context. Buy the full cycle around each error, not a clipped error moment, so models learn the deviation relative to correct execution. Clipping practices are covered in curating raw video into training clips.
- Reserve errors for evaluation. Hold out some rare error types entirely from training to test generalization to unseen mistakes. See long-tail and edge-case coverage for measurement methods.
Label errors after the fact by joining video to quality records
The most reliable way to find natural errors in hours of footage is to start from records that already say an error happened. Manufacturing and field-service operations log nonconformance reports (NCRs), rework tickets, scrap codes, failed inspection results, warranty returns and reopened work orders, each with a timestamp, station or asset, and often a defect code.
The join works like this: match the record's unit serial, work order or asset ID to the recording's station and time window; then have annotators locate the moment the defect was introduced and, where present, the correction. Defect codes map to your error types with a lookup table, which also exposes codes that do not fit and need a taxonomy decision. Owner page context on these records is on manufacturing quality datasets for AI training, and the agent-side equivalent is covered in rework, reversals and reopened cases.
Two failure modes to check. First, records catch errors that escaped to inspection; errors corrected on the spot leave no record, so corrections and near-misses still need direct review of footage. Second, timestamps in MES or ticketing systems are often entry times, not event times, so allow a search window and record its width.
Illustrative example: invented to show structure; it does not describe an available dataset.
{"clip_id":"st07_2026-03-14_cam2_0412","station":"ST07","task":"pump_housing_assembly","sop_version":"r12","view":"overhead_rgb","fps":30,"start_ts":"2026-03-14T10:41:07Z","end_ts":"2026-03-14T10:47:55Z","events":[{"type":"step","step_id":"S04","label":"seat_o_ring","t_start":112.4,"t_end":121.9},{"type":"error","error_class":"wrong_part","step_id":"S06","detail":"M6x20 used where SOP specifies M6x16","t_start":188.0,"t_end":196.3,"evidence":"ncr","record_ref":"NCR-0000-illustrative"},{"type":"correction","resolves":"S06","t_start":241.5,"t_end":259.0}],"label_source":"record_join+human_review","annotator_agreement":0.82,"pii_treatment":"faces_blurred;badge_ids_masked"}
Delivering one event-bearing record per line in JSON Lines keeps clips, steps and errors streamable and diffable [3]. Pair the timeline with the SOP version, since a "wrong order" label is only meaningful against the procedure in force that day; see pairing procedure video with written SOPs and temporal action segmentation labels.
Label quality checks specific to error classes
Error labels are noisier than step labels, and the noise concentrates in the classes you most care about. Benchmark audits have found average test-set label error rates of at least 3.3% across widely used datasets, enough to change which model looks best [2]; on a class with a few dozen examples, two or three wrong labels can swing measured recall substantially.
Ask suppliers or annotation vendors for:
- Double annotation on every error and correction segment, with agreement reported per error class rather than overall.
- An adjudication log for disagreements, including taxonomy changes made mid-project.
- The share of error labels confirmed by a quality record versus found by review alone.
- Boundary rules: when an omission "starts" (the moment the step was due) and when a technique error ends.
Supplier sensitivity: error footage carries liability, so expect stricter approval
Companies are more cautious releasing footage of mistakes than footage of correct work, because it can show defects, safety lapses, or identifiable workers making errors. Expect the supplier's quality, legal and HR functions to review requests, and expect some to decline outright or to restrict clips involving injuries, regulated products or customer-identifiable units.
What helps a request get approved: ask for de-identified workers (faces blurred, badges and name tags masked, audio removed or transcribed and redacted), strip customer names, serials and order numbers from frames and joined records, and specify allowed uses clearly. Worker notice and recording rules are covered in recording employees on video for AI datasets, and the full stack of rights in a single clip in rights layers in a video clip.
Buyer checklist for mistake detection video
Illustrative example: invented to show structure; it does not describe an available dataset.
| Question | Why it matters | Acceptable evidence |
|---|---|---|
| How did errors occur (natural, induced, scripted)? | Determines visual realism of errors | Collection protocol, share of each type per class |
| Which error taxonomy, and is it versioned? | Prevents class drift across batches | Taxonomy document with definitions and examples |
| How many error events per class, per station? | Rare classes may be untrainable | Per-class counts with normal-cycle denominators |
| Are corrections and near-misses labeled? | Needed for intervention timing | Sample timelines showing correction links |
| What records were joined, and how? | Validates labels and timing | Join keys, time window width, unmatched rate |
| What SOP version applies to each clip? | "Wrong order" depends on procedure | SOP revision per clip |
| How were workers de-identified? | Privacy and supplier approval | Method description and a checked sample |
| Which views and sensors? | Egocentric vs overhead changes what is visible | Camera layout, fps, resolution, sync method |
Egocentric capture of skilled manual work, which shows hands and tools at close range, is discussed on egocentric video datasets of skilled manual work and in manufacturing assembly video datasets.
How SourceX approaches error-bearing procedural video
SourceX sources operational data from US companies on request, including new recordings of hands-on work, and manages the licensing and ongoing purchases. Nothing is held in stock and a request does not guarantee a match; you describe the data you need, such as error types, task domain, views and record joins, and SourceX looks for businesses that hold it. Each dataset is rights-reviewed, personal details are removed or replaced before delivery with the method recorded and a sample checked, and every release is approved by the supplying company. You can start a buyer request once your taxonomy and volume targets are drafted.
Request mistake detection video data
If your task-verification or quality-assist model needs real errors, corrections and near-misses rather than staged ones, describe the tasks, error classes and labels you need. SourceX works through Find, Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Describe the task errors you need on video.
Sources
- Sener et al. (CVPR 2022), arXiv, "Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities" (2022). https://arxiv.org/pdf/2203.14712
- Northcutt, Athalye, Mueller (NeurIPS 2021), arXiv, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- jsonlines.org, "JSON Lines". https://jsonlines.org/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.