Skip to content

Multimodal and embodied data

Warehouse Picking and Packing Data for Robot Manipulation Models

Quick answer

A useful robotic picking dataset is not a folder of bin photos. It is a set of real pick attempts in which each station image or depth frame is joined to the item master record (dimensions, weight, packaging type) and to a verified outcome: success, drop, double pick, damage or no-grasp. Public sets show the shape of this data, but production-grade coverage of long-tail SKUs, deformable bags and transparent packaging usually has to be licensed from operators whose warehouse systems hold those joins.

By SourceX Editorial · Updated

What a pick-level record has to contain

A pick-level record is useful for grasp prediction only when perception, item identity and outcome share one key, usually a pick task ID from the warehouse management system (WMS) or warehouse execution system (WES). The image alone tells a model what the tote looked like; the item master tells it what it was trying to grasp; the outcome tells it whether the attempt worked. Drop any one of the three and you are left with a segmentation set or a logistics table, not manipulation training data.

The public reference point is ARMBench, recorded in an Amazon robotic workcell where a manipulator singulates items from mixed-content containers. It captures images, videos and metadata at pre-pick, transfer and post-placement stages for more than 235,000 pick-and-place activities on more than 190,000 unique objects [1]. Its benchmark tasks are segmentation in clutter, object identification and defect detection [1], which shows what a well-joined pick record supports and also what it leaves to the buyer: gripper type, suction parameters and your own end effector are not in someone else's workcell logs.

Typical joins a buyer should ask about:

  • Station sensors: RGB frames, depth or point clouds (often as 16-bit PNG depth or PLY), camera intrinsics and extrinsics, and timestamps aligned to the robot controller.
  • Item master (from the WMS/ERP): SKU, GTIN or UPC, length, width, height, weight, packaging type (polybag, carton, blister, bottle), fragility and hazmat flags.
  • Task context: source tote or bin ID, destination, order line, units requested versus picked.
  • Outcome: grasp result codes, retry count, exception codes from the WES, and downstream signals such as weight-check failures or pack-station rejects.

For contact signals beyond these fields, see our page on tactile and force-torque data for contact-rich manipulation.

Why outcome labels from warehouse systems need auditing

Outcome fields in warehouse systems are operational codes, not ground truth, so they need validation before you train on them. A "pick complete" event may be written when a scan succeeds, even if the item was damaged in transfer, and a drop recovered by a human may be logged as a normal retry. Check how each code is generated and compare a sample against video.

Pick logs also carry a selective-labels problem: you only observe outcomes for grasps the deployed policy chose to attempt [3]. Items the system routed to manual picking, or grasp points it never tried, have no outcome at all, which biases a learned success predictor toward easy items. Ask the supplier for the routing rules and for manual-exception volumes per SKU class, and read our guide to verifying outcome labels in operational records.

Covering the long tail of SKUs and packaging

Long-tail coverage is the main reason to license real warehouse data rather than rely on public or simulated sets. The failure cases that matter in production are concentrated in deformable polybags, shrink-wrapped multipacks, transparent and reflective items that break depth sensing, and mixed totes where items overlap. A dataset dominated by rigid cartons will look strong on aggregate metrics and still fail on those slices.

Specify coverage by packaging type and by optical property, not by total image count. Ask for per-slice counts of attempts and failures, the number of distinct SKUs per slice, and how many sites and camera rigs contributed. Pooled academic data shows the value of breadth across embodiments: Open X-Embodiment combines over one million trajectories from 22 robot embodiments [2], but it is not drawn from the e-commerce assortment of a working fulfillment center. For the tradeoff with synthetic data, compare real versus simulated robot data.

Illustrative example: invented to show structure; it does not describe an available dataset.

Coverage sliceWhat to requestCommon failure if missing
Polybags and pouchesAttempts and failures per SKU, suction seal outcomeSuction leaks and slips
Transparent or reflective itemsRaw depth with invalid-pixel masks, not filled depthMissing or wrong grasp points
Mixed, cluttered totesInstance masks and fill level per toteDouble picks, collisions
Heavy or oversize itemsWeight from item master, payload faultsDrops in transfer
New SKUs (first 30 days)Attempts before item master was completeOverstated performance on known items

Rights in third-party logistics data

In third-party logistics, the warehouse operator usually does not own everything in its own pick records. Product images, labels, item master attributes and order data may belong to the 3PL's clients under fulfillment contracts, while the robot OEM may claim rights in controller logs. Before negotiating, map who holds rights in each layer and whether client contracts allow use of their product data for model training.

The same question applies to vendor data: some robot cell vendors restrict how operators share logs. Our guides on robot data ownership across operators, OEMs and facility owners and on licensing multimodal records from several rightsholders cover how to structure those permissions. For the supply-chain side, see supply chain and logistics datasets and third-party logistics buyers.

People in picking and packing station footage

Station footage often captures human pickers, packers and maintenance staff, so treat it as personal data. Faces and hands appear at manual pick cells, induction stations and pack benches, and the consent rules differ by state. Illinois BIPA covers scans of face geometry and requires a retention schedule and a written release, which includes electronic signatures [4], and Texas Business and Commerce Code Section 503.001 requires notice and consent before capturing a biometric identifier for a commercial purpose [5].

Ask suppliers how they handle face blurring, badge and screen redaction, and audio stripping, and whether employee notices covered recording for model training. Activity-recognition sets such as DaRA, which recorded order picking and packaging in a lab built to resemble a warehouse [7], show how much human motion even a controlled setup captures. For methods, read de-identifying multimodal records.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Request template for a robotic picking dataset

A specific request gets better answers than a category name. Describe the station, the joins and the outcomes, and ask for documentation in a structured form such as a data card covering sources, collection method, annotation and intended use [6].

Illustrative example: invented to show structure; it does not describe an available dataset.

request: robotic picking data, e-commerce fulfillment
stations: goods-to-person pick cells, pack stations
sensors: [rgb, depth_or_pointcloud, camera_calibration]
join_key: wes_pick_task_id
item_master_fields: [sku, gtin, dims_mm, weight_g, packaging_type, fragile_flag]
outcome_fields: [grasp_result, retry_count, exception_code, weight_check_result]
coverage_targets:
  polybag_share: ">= 25% of attempts"
  transparent_or_reflective: "reported separately"
  failures: "kept, not filtered out"
people_in_frame: blurred faces; no audio
documentation: data card, field dictionary, outcome code definitions
intended_use: grasp prediction training, held-out evaluation

Before accepting delivery, run the checks in acceptance checks for robot datasets: join completeness, timestamp alignment between camera and controller, and the share of failures retained.

How SourceX approaches pick and pack data requests

SourceX sources operational datasets from US companies and manages the licensing process, including agreements and ongoing purchases. Picking data is sourced on request and is not held in stock, so a request describes the data you need rather than a business, and it does not guarantee a match. SourceX looks for US businesses that hold the described data, and every release is approved by the supplying company; you can describe your picking data need to SourceX.

The process runs Find, Assess (the data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Names and other personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect.

Delivery runs through private, access-controlled workflows after an executed agreement and supplier approval. SourceX does not source generic CCTV or photos and does not train models. For the wider category, see robotics training data for embodied AI, the multimodal and embodied data hub and the AI data guides.

Sourcing real robotic picking data

If your grasp or packing models need real pick attempts joined to item master data and outcomes, start with a written description of stations, fields and coverage. SourceX sources operational datasets on request from US companies and handles rights review and licensing. Describe your robotic picking dataset request.

Sources

  1. Papers with Code (Mitash et al., Amazon), "ARMBench: An Object-centric Benchmark Dataset for Robotic Manipulation" (2023). https://astro.paperswithcode.com/dataset/armbench
  2. arXiv (Open X-Embodiment Collaboration), "Open X-Embodiment: Robotic Learning Datasets and RT-X Models" (2023). https://arxiv.org/abs/2310.08864v1
  3. KDD 2017 (Lakkaraju, Kleinberg, Leskovec, Ludwig, Mullainathan), "The Selective Labels Problem: Evaluating Algorithmic Predictions in the Presence of Unobservables" (2017). https://www.cs.cornell.edu/home/kleinber/kdd17-selective.pdf
  4. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  5. Texas Legislature, "Texas Business and Commerce Code Section 503.001, Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  6. arXiv / FAccT 2022 (Pushkarna, Zaldivar, Kjartansson), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  7. PubMed Central, "DaRA Dataset: Combining Wearable Sensors, Location Tracking, and Process Knowledge for Enhanced Human Activity Recognition". https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12846308/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data