Skip to content

Video data

Manufacturing Assembly Video Datasets for Action Recognition

Quick answer

A useful assembly video dataset is real production footage from fixed or wearable cameras at assembly stations, labeled with operations, parts, tools and cycle boundaries that line up with the plant's work instructions and MES timestamps. Public benchmarks such as IKEA ASM and Assembly101 are staged recordings of furniture or toy kits, so they rarely transfer to torque sequences, harness routing or fixture changes. Commercial use needs a license from the manufacturer, plus worker-privacy and design-confidentiality review before any frame leaves the plant.

By SourceX Editorial · Updated

Why lab assembly benchmarks fall short for production models

Lab benchmarks are good for pretraining and method comparison, but they miss the variance that breaks models on a real line. Assembly101, for example, records multi-view assembly and disassembly of toy vehicles for procedural activity research [1], and other academic sets use furniture kits or toy models assembled by volunteers.

Production lines differ in ways that matter to the model. Operators are trained and fast, so a "fasten" step can last as little as a second in high-speed operations, and gloves hide hand pose. Parts are near-identical variants told apart by a label or connector color, and occlusion by fixtures, conveyors and the product itself is constant. Check each benchmark's own license file before any commercial use, since research datasets often restrict it, and treat them as a starting point rather than a substitute (see the video data hub for how sources compare).

Labels that make assembly footage trainable

The labels that matter tie each video segment to a step the plant already defines, not to a generic verb list. Build the taxonomy from the work instruction (WI) or standard work sheet: one label per WI step, with part number, tool ID and expected duration. Then align cycle boundaries with MES or PLC events such as a station-start barcode scan, a torque tool's OK/NOK result or a pallet-release signal.

That alignment gives you weak labels for free and a check on human annotation. If the annotated "fasten bolt 3" segment does not overlap a logged torque event, one of them is wrong. Our guide to temporal action segmentation labels covers boundary tolerance and inter-annotator agreement; pairing footage with written procedures is covered in video-to-SOP step alignment.

Defect and deviation labels are the scarce part. Skipped steps, wrong-order steps, rework loops and missing fasteners are rare by design, so ask how many cycles end in a NOK or rework flag before you commit. Quality outcome records, covered on our manufacturing quality datasets page, are how you find those cycles.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "clip_id": "stn07_cam_overhead_2026-03-14T09-12-04Z",
  "station": "ST-07 door module",
  "camera": {"view": "overhead_fixed", "fps": 30, "resolution": "1920x1080"},
  "cycle": {"serial_hash": "a91f...", "mes_start": "09:12:04.210", "mes_end": "09:13:01.880", "result": "rework"},
  "segments": [
    {"start_s": 0.0, "end_s": 4.3, "wi_step": "WI-0712-03", "action": "pick", "part": "BRKT-L", "tool": null},
    {"start_s": 4.3, "end_s": 9.8, "wi_step": "WI-0712-04", "action": "fasten", "part": "M6 bolt x3", "tool": "DC-nutrunner-02", "torque_events": 2},
    {"start_s": 9.8, "end_s": 11.1, "wi_step": null, "action": "deviation", "note": "third bolt skipped"}
  ],
  "privacy": {"faces": "masked", "badges": "masked", "method_log": "redaction_v3"},
  "confidentiality": {"supplier_approved": true, "masked_regions": ["unreleased part label"]}
}

Camera placement decides what you can label

Camera position fixes the label ceiling, so specify it before anything else. An overhead station-fixed camera sees part placement and sequence well and supports cycle-time analytics, but loses fine finger actions. Side views catch reach, insertion depth and posture. Wearable cameras capture hand-object contact for work-instruction and robot learning models, but frame the operator's environment unpredictably; our page on egocentric video of skilled manual work covers that format.

Ask for synchronized multi-view where possible. A second view can recover steps hidden by fixtures or the operator's body, and multi-view benchmarks such as Assembly101 record simultaneous views of each session [1]. Teams training manipulation policies should also read human demonstration video for robot learning.

Worker privacy without destroying the signal

Masking faces is usually cheap for model quality, but masking whole bodies is not. In a study of detection and pose training, traditional anonymization noticeably hurt performance, full-body anonymization cost more than face-only, and realistic generative anonymization reduced the loss [2]. That work used still images, so verify the effect for your own action task before fixing a redaction policy.

For assembly footage, face and badge masking usually preserves the hands and tools the model needs. Biometric law still applies: Illinois BIPA covers scans of face or hand geometry and requires written release (including electronic signatures) and a retention schedule; 2024 amendments limit liability to a single recovery per person for multiple collections [3]. Notice, consent and audio rules for recording staff are covered in recording workers on video for AI datasets, and works-council or union agreements may also govern recording at some sites. For redaction methods generally, see the privacy hub.

Design confidentiality and supplier approval

Assembly footage shows things the manufacturer considers secret: unreleased parts, fixtures, line layout and takt times. Expect the supplier to review sample frames, mask part labels or CAD-derived features, and exclude certain stations. Plan for that early, because a masked region over the part being assembled can make a clip useless.

Separate the rights layers too: the footage, the people in it, visible displays or HMIs, and any third-party branded components. Rights layers in a video clip breaks these down.

Buyer checklist for an assembly video request

Describe the data, not a named company, and make each requirement testable. Use this as a request template.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldWhat to specifyWhy it matters
Process typeManual, semi-automated or cobot-assisted; product family (for example wire harness, electronics, door modules)Action vocabulary and speed differ by process
ViewsOverhead fixed, side, wearable; synchronized or not; fps and resolutionSets the label ceiling
Volume unitStation-hours or complete cycles, per variantClip counts hide repetition
LabelsWI-step taxonomy, part, tool, cycle start/end, deviationsNeeded for segmentation and verification
System alignmentMES cycle IDs, torque OK/NOK, barcode scans, clock sync methodWeak labels and QA checks
Rare eventsShare of cycles with rework or NOKDefect detectors need positives
PrivacyFace and badge masking, audio removed, method logWorker protection, BIPA exposure
ConfidentialityMasked regions, excluded stations, approval stepAvoids late rejections
Allowed usesTraining, evaluation, fine-tuning, model releaseMust match your roadmap
DeliveryContainer format, sharding, manifest with checksumsMulti-terabyte handling

For delivery, ask for shards plus a manifest rather than loose files; WebDataset tar shards, often about 1 GB each, are a common pattern for datasets that reach multiple terabytes [4]. Our delivery formats guide covers manifests and transfer. If you also need machine signals and quality outcomes in sync with the video, see synchronized manufacturing data; for governing data quality, including labels, across the data life cycle, ISO/IEC 5259-5 offers a framework [5].

How SourceX handles assembly video requests

SourceX sources operational datasets from US companies, including new recordings of hands-on work, and manages the commercial process through licensing and ongoing purchases. Data is sourced on request rather than held in stock, so a request does not guarantee a match, and SourceX does not source generic CCTV footage. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and every release is approved by the supplying company. Buyers can read how the process works on the buyers page, and manufacturing context is on buyers by industry: manufacturing. For background on demand, see do AI labs buy video of people working?

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Request production assembly video for your models

Describe the stations, views, labels and allowed uses you need, and SourceX looks for US businesses that hold matching footage. Nothing is contracted until a supplier agrees, and delivery runs through private, access-controlled workflows under a license that defines records, uses, term and delivery. Describe the assembly video you need.

Sources

  1. Sener et al., CVPR 2022 (CVF Open Access), "Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities" (2022). https://openaccess.thecvf.com/content/CVPR2022/html/Sener_Assembly101_A_Large-Scale_Multi-View_Video_Dataset_for_Understanding_Procedural_Activities_CVPR_2022_paper
  2. Hukkelas and Lindseth, CVPR 2023 Workshops (CVF Open Access), "Does Image Anonymization Impact Computer Vision Training?" (2023). https://openaccess.thecvf.com/content/CVPR2023W/WAD/papers/Hukkelas_Does_Image_Anonymization_Impact_Computer_Vision_Training_CVPRW_2023_paper.pdf
  3. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)" (current text). https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  4. Hugging Face, "WebDataset (Hub documentation)". https://huggingface.co/docs/hub/datasets-webdataset
  5. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-5:2025 Data quality for analytics and machine learning, Part 5: Data quality governance framework" (2025). https://www.iso.org/standard/5259-5

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data