Image data
Retail Shelf Image Datasets for SKU Detection and Planogram Compliance
Quick answer
A retail shelf image dataset for production models needs more than dense boxes. Public benchmarks such as SKU-110K teach a detector where products are, but its annotations carry no SKU categories, so recognition, out-of-stock and planogram compliance models need store photos linked to SKU labels, planograms and audit outcomes [1][2]. Buyers usually license that layer from retail execution, merchandising or store-audit operations and specify store formats, fixtures, cameras and label taxonomy before sourcing.
By SourceX Editorial · Updated
What public shelf benchmarks cover and where they stop
Public shelf benchmarks cover dense localization well and SKU identity poorly. SKU-110K, the standard dense-detection benchmark, ships 8,219 training, 588 validation and 2,936 test images of packed shelves [1]. Its labels are class-agnostic boxes, which one retail recognition paper notes rules it out for SKU recognition on its own [2].
Newer sets narrow the gap but not for your assortment. Retail-YU pairs shelf and web images for 1,505 SKUs across roughly 103,000 images [4], which is useful for pretraining a retrieval embedding, but its SKU list will not match a specific retailer's planogram. Teams publishing on retail detection often fall back on in-house data: one semi-supervised study used about 9,000 proprietary shelf images averaging around 102 products each [3].
The practical reading: use public sets to pretrain the detector head and benchmark density handling, then license operational images for the labels that carry commercial value. For pairing shelf crops with catalog packshots, see SKU recognition training data with shelf and catalog image pairs.
Label layers a shelf monitoring dataset needs
A shelf monitoring dataset needs at least four label layers: product boxes, SKU identity, shelf structure and a compliance outcome. Boxes alone train a facing counter; the other layers turn it into a planogram or out-of-stock model.
- Facings: tight axis-aligned boxes per visible product face, COCO JSON with
bboxas[x, y, width, height]in pixels so it round-trips through CVAT and common trainers [5]. Standard COCO boxes cannot express rotation, so tilted products need polygons or an agreed extension. - SKU identity: a GTIN/UPC or internal item number per box, plus an "unknown" class for items absent from the master file. Without this, recognition becomes a separate retrieval project.
- Shelf geometry: shelf-edge lines or bay and shelf indices, so detections map to planogram positions (bay 3, shelf 2, position 7).
- Gap and condition labels: empty slot, partial facing, price-tag mismatch, misplaced item, toppled or unfaced products not pulled forward.
- Audit outcome: the image-level or slot-level result the field rep recorded, such as compliant, out-of-stock or misplaced, with the audit timestamp.
The richest source of the last two layers is not an annotation vendor but existing store-audit and merchandising workflows, where reps already photograph bays and tick outcomes against a planogram. Those records double as weak labels and as an evaluation set for your annotation and quality decisions.
Capture variables that decide whether a shelf model transfers
Shelf models fail most often on capture shift, not on class count. Specify the capture context explicitly, because a detector trained on clean handheld photos from one grocer degrades on fisheye ceiling cameras or robot passes in a pharmacy.
Write these into the request: store formats (grocery, convenience, drug, mass, club, specialty), fixture types (gondola, cooler doors with glare, end caps, pegboards, checkout racks), camera source (rep handheld, autonomous robot, fixed shelf or ceiling camera), resolution and stitching (single bay vs panorama), lighting, and time-of-day restock cycles. Per-image density matters too; a set dominated by sparse end caps will not train the dense-scene behavior that benchmarks like SKU-110K stress [1][3]. For more on sensor and lighting spread, see camera, lens and lighting diversity in image datasets.
Split by store and by time, not by image. Random image splits leak the same bay across train and test, which inflates mAP and hides planogram-reset failures.
Shelf image request template
A useful request names the task, the label layers and the capture mix in one page so a data holder can confirm what it has.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example specification |
|---|---|
| Task | Slot-level out-of-stock and planogram compliance; facing count; SKU recognition for top 2,000 items |
| Store formats | Grocery and drug, US, mix of urban and suburban |
| Fixtures | Ambient gondolas, cooler doors, end caps |
| Camera source | Rep handheld (majority), fixed shelf camera (minority) |
| Images | Bay-level JPEG, original resolution, EXIF capture time kept, GPS removed |
| Labels | COCO boxes, GTIN per box, shelf index, gap class, audit outcome |
| Linked records | Planogram version per bay (position, SKU, facings), audit form result |
| Splits | Held-out stores and held-out months |
| People in frame | Faces and bodies blurred before delivery; blur method documented |
| Documentation | Data card covering source, capture, annotation and known gaps [7] |
A matching record might look like: image_id, store_format: grocery, fixture: cooler_door, camera: handheld, planogram_id, bay: 4, annotations: [{bbox, gtin, shelf: 2, position: 6, status: out_of_stock}].
Rights and privacy issues specific to store photos
Store shelf photos carry three rights questions: who owns the photo, what the retailer permitted, and who appears in frame. Merchandising agencies and brand field teams often capture photos under a retailer's store-access or photo policy, so the supplying company must confirm it can license those images for model training, not just for reporting.
Packaging shows brand trademarks and artwork, which matters less for detection training than for anything that reproduces images; still, record the permitted use in the license. Shoppers and staff caught in aisle shots should be blurred, and buyers should avoid any face-derived signal: the FTC's action against Rite Aid's in-store facial recognition shows the regulatory exposure when store imagery drifts into identifying people [6]. See licensing image and video data that contains faces and model and property releases for AI training images.
Also check EXIF: GPS coordinates identify store locations, which can be competitively sensitive for the supplier. The EXIF metadata guide covers what to strip and keep.
How SourceX sources shelf imagery for buyers
SourceX sources operational datasets from US companies on request; shelf imagery is not held in stock, and a request does not guarantee a match. Buyers describe the data, such as store formats, fixtures, cameras and label layers, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. It does not source scraped web content or generic CCTV or photos.
Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, and the process runs Find, Assess, Agree, Transact and Manage, with nothing contracted until a supplier agrees. Start at the SourceX buyer page, or browse the image data hub, images and inspection photos, product catalogs and descriptions and multi-brand retail buyers.
Source retail shelf images for your planogram models
SourceX serves AI teams wherever they are based and manages the commercial process from licensing to ongoing purchases. Describe the shelf formats, cameras and label layers you need, and pricing and allowed uses are agreed per deal in a license. Describe the shelf image data you need.
Sources
- Ultralytics, "SKU-110K Dataset". https://docs.ultralytics.com/datasets/detect/sku-110k
- arXiv, "RP2K: A Large-Scale Retail Product Dataset for Fine-Grained Image Classification" (2020). https://arxiv.org/pdf/2006.12634v1
- arXiv, "Semi-supervised Learning for Dense Object Detection in Retail Scenes" (2021). https://arxiv.org/pdf/2107.02114
- Mendeley Data, "Retail-YU shelf and web product image dataset" (2025). https://data.mendeley.com/datasets/mmcf24t9vv/1
- CVAT.ai, "COCO (CVAT format documentation)". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
- Federal Trade Commission, "Coming face to face with Rite Aid's allegedly unfair use of facial recognition technology" (2023). https://www.ftc.gov/business-guidance/blog/2023/12/coming-face-face-rite-aids-allegedly-unfair-use-facial-recognition-technology
- Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.