Video data
Warehouse Operations Video Datasets: Picking, Packing, Loading and Forklift Footage
Quick answer
A useful warehouse video dataset is task-focused footage of real picking, packing, palletizing, dock loading and forklift travel, recorded from cameras placed to see hands, totes and forks, and joined to warehouse management system (WMS) events that act as cheap weak labels. Buy it with per-clip task labels, site-level metadata, worker and vehicle anonymization, and a license that names your model uses. Generic security CCTV rarely meets that bar.
By SourceX Editorial · Updated
What public warehouse video covers, and where it stops
Public warehouse video is mostly lab-staged or robot-only, so production-floor footage usually has to be licensed from operators. Academic activity-recognition sets often record order picking and packing in labs built to resemble a warehouse, with actors following scripts. Robotic pick-and-place benchmarks capture grasps and transfers but contain no human operators, dock doors or forklifts.
General egocentric corpora such as Ego4D show what consented first-person collection looks like at scale, yet warehouse tasks are a thin slice of them [1]. Robot learning teams also face fragmented demonstration data; Open X-Embodiment had to pool many separate robot datasets into one format to train across embodiments [2]. That gap is why buyers turn to real operational footage from distribution centers, third-party logistics sites and cross-docks. For the robot-side view, see warehouse picking and packing data for robot manipulation.
A task taxonomy buyers can put in a request
Define the label space before you ask for footage, because the taxonomy decides camera placement, clip length and price. A working taxonomy separates operator tasks, vehicle tasks and exceptions, and records the WMS event that confirms each one.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Task class | Typical camera | Weak-label source | Hard negatives and edge cases |
|---|---|---|---|
| Each pick (shelf, flow rack) | Overhead at pick face, chest-mounted | Pick confirmation scan, RF gun timestamp | Short picks, wrong-slot picks, reach into adjacent bin |
| Pack and cartonize | Fixed overhead at pack station | Pack-out scan, carton label print | Dunnage only, re-pack, multi-order totes |
| Scan and verify | Station or wrist camera | Barcode read events | Failed reads, manual key entry |
| Palletize and stretch-wrap | Elevated side view | Pallet build complete (LPN) | Mixed-SKU layers, collapsed tiers |
| Trailer load and unload | Dock door camera | Ship confirm, ASN receipt | Floor-loaded cartons, partial loads |
| Forklift travel and put-away | Mast or ceiling camera, onboard camera | Put-away confirmation, telematics | Reversing, pedestrian crossings, near-misses |
| Exceptions | Any of the above | Short-pick, damage or hold codes | Rare classes; plan oversampling |
Forklift footage deserves its own spec. OSHA.s powered industrial truck standard, 29 CFR 1910.178 [7], covers counterbalance forklifts, reach trucks, order pickers and powered pallet jacks, and its training and operating rules give safety models a defensible vocabulary for unsafe acts. Near-miss and incident clips are covered in workplace safety video data.
Turning WMS events into weak labels
Joining video to WMS events is the cheapest way to label warehouse footage at scale. Each pick confirmation, scan or put-away carries a timestamp, user ID, location and SKU, so a model team can cut a window around the event and treat it as a candidate segment for "pick" or "put-away." Human annotators then refine boundaries rather than search raw hours.
Three failure modes recur. Camera clocks drift against WMS server time, so ask for an NTP sync record or a visible sync marker per shift. Batch confirmations, where a worker scans after several picks, smear event timing, so flag batch-mode workflows. Labor management system (LMS) user IDs are personal data and must be pseudonymized consistently across video and logs. The underlying event data is covered in WMS task, exception and labor logs, and step-level labeling conventions in temporal action segmentation labels.
Splits, quality and delivery format
Hold out whole videos, shifts or sites rather than random frames, because fixed warehouse cameras repeat near-identical scenes and frame-level splits inflate scores. A defensible split puts at least one site, or one camera layout, entirely in test. Ask for per-clip metadata: site ID (pseudonymous), camera ID, mount height and angle, resolution, frame rate, lighting, shift, and the WMS or LMS system family.
Label quality needs its own evidence. ISO/IEC 5259-4 frames data quality processes for ML training, including labelling of supervised data, and is a reasonable reference for asking how agreement was measured [4]. For packaging, bounding boxes and tracks commonly travel as COCO JSON from tools such as CVAT [5], while large clip sets stream well as WebDataset tar shards [6]. The video dataset packaging guide covers sidecar metadata and clip indexes.
Privacy, worker consent and rights checks
Warehouse video almost always shows identifiable workers, so anonymization and consent records belong in the deal, not after it. Faces, badges, name tags, forklift unit numbers, trailer and license plates, and screens on RF guns or pack stations can all identify a person or a client. Illinois's BIPA treats face-geometry scans as biometric identifiers, so face-recognition derivatives raise obligations under BIPA [3], which was amended in 2024 to treat repeated collections of the same identifier as a single violation.
Ask the supplier how workers were notified, whether union or works-council agreements address recording, and whether audio was captured. Customer brands on cartons and labels add another rights layer. Our pages on recording employees for AI datasets, video anonymization and rights layers in a video clip go deeper.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Buyer checklist for a warehouse video request
A request that names tasks, cameras, labels and uses gets a clearer yes or no from suppliers. Use this as a starting template.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Tasks: each pick, pack, palletize, trailer load, forklift put-away; list exceptions you need.
- Cameras: task-focused mounts (pick face, pack station, dock door, mast); no general security CCTV.
- Volume and diversity: hours per task, number of sites, shifts and lighting conditions.
- Labels: WMS-joined weak labels plus human-verified segments on a stated share; agreement metric.
- Metadata: camera pose, fps, resolution, clock sync method, pseudonymous site and user IDs.
- Privacy: face, badge, plate and screen redaction method and sample QA result.
- Rights: worker notice records, customer-brand handling, allowed uses (training, evaluation, fine-tuning).
- Delivery: container, annotation format, shard layout, holdout definition.
How SourceX handles warehouse video requests
SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and manages the licensing process; it does not hold warehouse footage in stock, and a request does not guarantee a match. SourceX does not source generic CCTV, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, with names and other personal details removed or replaced before delivery. You can describe the warehouse footage you need to SourceX, or start from the video data hub, the logistics buyer page, third-party logistics, robotics and embodied AI data and video recordings.
Sourcing warehouse operations video for your model
- Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)" (as amended 2024). https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004 If you need real picking, packing, loading or forklift footage with labels and a commercial license, describe the tasks, cameras and uses rather than naming a company. SourceX looks for US businesses that hold that data, reviews ownership and consents, and agrees pricing and allowed uses in a license before anything is delivered. Tell SourceX what warehouse video you need.
Sources
- arXiv, "Ego4D: Around the World in 3,000 Hours of Egocentric Video" (2021). https://arxiv.org/pdf/2110.07058
- arXiv, "Open X-Embodiment: Robotic Learning Datasets and RT-X Models" (2023). https://arxiv.org/pdf/2310.08864
- Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
- ISO/IEC, "ISO/IEC 5259-4:2024 Data quality process framework" (2024). https://www.iso.org/standard/81093.html
- CVAT documentation, "COCO format". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
- GitHub (webdataset project), "webdataset". https://github.com/webdataset/webdataset
- Occupational Safety and Health Administration (OSHA), "29 CFR 1910.178 - Powered industrial trucks". https://www.ecfr.gov/current/title-29/subtitle-B/chapter-XVII/part-1910/subpart-N/section-1910.178
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.