Image data
SKU Recognition Training Data: Pairing Shelf Crops with Catalog Reference Images
Quick answer
SKU recognition needs different data from shelf detection. A detect-then-recognize pipeline needs box-level shelf annotations to find products, then an identity layer: a reference gallery of catalog images per SKU, plus shelf crops labeled with the same SKU IDs so the embedding model learns to bridge the studio-to-store domain gap. The hardest requirements are coverage of near-duplicate variants (pack size, flavor, promo packs), dated packaging versions, and clear rights to train on brand packaging.
By SourceX Editorial · Updated
Why recognition data is separate from detection data
Recognition data answers "which SKU is this," while detection data only answers "where is a product." Many retail recognition systems split the problem: a class-agnostic or coarse detector proposes boxes on the shelf, and a second model identifies the SKU from each crop [1]. Detection benchmarks such as SKU-110K are built for the first stage, with densely packed, look-alike items labeled as boxes rather than as thousands of distinct identities [3].
That split matters for procurement. Box labeling in dense scenes is expensive enough that researchers turn to semi-supervised detection to reduce it [4], but a box without a SKU ID does nothing for the recognizer. If you are sourcing detection data, start with our guide to retail shelf image datasets for SKU detection and planogram compliance; this page covers the identity layer.
The recognizer also changes more often than the detector. Assortments change constantly as SKUs launch and are delisted, so many teams train a metric-learning embedding (ArcFace, triplet or supervised contrastive loss) and recognize by nearest-neighbor lookup against a gallery. Adding a SKU then means adding reference images, not retraining a 10,000-way classifier.
What a reference gallery per SKU should contain
A reference gallery is the set of clean, canonical images that defines each SKU's identity at inference time. The minimum is one front-facing catalog image per SKU, which is exactly the "Web domain" design used in the dual-domain Retail-YU dataset: in-store shelf photos paired with one curated online image for each of 1,505 SKUs [2]. One image is enough to benchmark the domain gap, but production galleries usually want more.
Ask for these views and fields per SKU:
- Front face at high resolution, plus side and top faces for products that are faced sideways or stacked (cans, cereal boxes, bottles shown at an angle).
- Packshot provenance: whether the image is a photographed product or a rendered 3D/flat artwork file. Renders lack glare, curvature and shrink-wrap reflections, which widens the gap to shelf crops.
- Stable identifiers: GTIN or UPC, internal item number, brand, sub-brand, variant, net content and unit of measure, so near-duplicates can be disambiguated programmatically.
- Taxonomy: a category path, ideally mapped to GS1 Global Product Classification at segment, family, class and brick level [5], which lets you build hierarchical losses and evaluate errors within the same brick.
- Version dates: the date each packshot became valid and, where available, the date it was superseded.
If you are buying catalog imagery mainly for attribute extraction or e-commerce search, the product catalog photo datasets guide covers that use; for licensed catalog content in general, see SourceX's page on product catalogs and descriptions.
How shelf crops close the shelf-to-catalog domain gap
Shelf crops labeled with the same SKU IDs as the gallery are the only reliable way to measure and close the domain gap. Catalog images are centered, evenly lit and unoccluded; shelf crops are small, blurred, partly hidden by price rails and shelf talkers, distorted by perspective and colored by store lighting and fridge doors. A model trained only on catalog images can look strong on held-out catalog images and still fail on real shelves.
Specify the shelf side of the pair as carefully as the gallery:
- Crop source: crops cut from full shelf images with the box coordinates retained, not pre-cropped thumbnails, so you can re-crop with different padding.
- Capture diversity: multiple stores, fixture types (ambient shelf, cooler door, endcap, peg hook), devices (handheld phone, shelf-scanning robot, fixed camera) and lighting.
- Label linkage: each crop's SKU ID must resolve to a gallery entry; crops that could not be identified should carry an explicit
unknownorambiguouslabel rather than being dropped, because open-set rejection is a real production task. - Hard negatives: crops of visually adjacent SKUs from the same shelf, which give metric-learning losses their most useful pairs.
Keep train, validation and test splits separated by store and by capture date, not by random crop. Crops of the same facing photographed minutes apart are near-duplicates and will inflate retrieval accuracy if they leak across splits.
Covering near-duplicate SKUs and packaging versions
Near-duplicate variants and packaging changes cause most recognition errors, so coverage of them should be an explicit requirement. The confusions that matter in practice are pack-size variants (a 12 oz and 16 oz bag with the same artwork), flavor or scent variants that differ by a color band or a small word, multipacks versus singles, limited-edition and seasonal wraps, and "new look" redesigns that run beside old stock for months.
Ask suppliers to report, per category, how many SKU groups share a brand and sub-brand and differ only in variant attributes, and how many shelf crops exist for each member of those groups. A gallery that covers 5,000 SKUs but has shelf crops for only the best-selling variant of each family will not teach the fine-grained boundary you care about. The same imbalance logic applies to rare defect coverage in inspection data: the long tail is where purchased data earns its cost.
Packaging versions need their own history. Request a SKU version table that records each artwork change with an effective date, and capture dates on every shelf image, so you can match a 2024 shelf crop to the 2024 packshot rather than the current one. Without it, your labels silently mix two visual identities under one ID.
Rights, logos and personal data in retail imagery
Product packaging is brand intellectual property, so training rights need to be confirmed for both the catalog images and the shelf photos. Catalog packshots are often supplied to retailers under syndication terms written for display on product pages, not for model training; confirm the permitted use with whoever holds the license. Image models can memorize and reproduce training images, including trademarked logos [6], which matters most if any generative use is in scope.
Shelf photos add other risks: shoppers, staff and name badges in frame, and store layouts that a retailer may treat as confidential. Ask how faces and people were handled, and whether the supplier had permission to share images captured in its stores. For a fuller treatment of releases, see model and property releases for AI training images, and for metadata such as GPS and device serials, EXIF metadata in image training data. Licensed shelf and field photography more broadly is covered on images and inspection photos.
On these points, SourceX rights-reviews each dataset for ownership and consents and delivers it under a license that defines the records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe your recognition data requirement to SourceX as a buyer.
Request template for a shelf-to-catalog SKU dataset
A structured request makes supplier answers comparable and exposes coverage gaps before you pay for data. Use the record schema below to describe what each paired item should look like, then attach your category scope and split rules.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"sku_id": "SKU-004417",
"gtin": "00012345678905",
"brand": "ExampleBrand",
"sub_brand": "Crunch",
"variant": "Sea Salt",
"net_content": "7.5 oz",
"gpc_brick": "Crisps/Chips (Shelf Stable)",
"package_version": { "artwork_id": "AW-2025-03", "valid_from": "2025-03-01", "valid_to": null },
"reference_images": [
{ "uri": "gallery/SKU-004417/front.png", "face": "front", "source": "photographed", "px": "2400x3000" },
{ "uri": "gallery/SKU-004417/side.png", "face": "side", "source": "render", "px": "1200x3000" }
],
"shelf_crops": [
{
"uri": "shelf/store-118/2025-06-14/img_0932.jpg",
"bbox_xywh": [1412, 806, 188, 241],
"capture_date": "2025-06-14",
"fixture": "ambient_shelf",
"device": "handheld_phone",
"occlusion": "price_rail_partial",
"label_status": "verified"
}
],
"confusable_skus": ["SKU-004418", "SKU-004402"],
"rights": { "catalog_training_use": "confirmed", "store_capture_permission": "confirmed" },
"pii_treatment": "faces_blurred; method_recorded"
}
Pair the schema with a short checklist for supplier responses:
| Question | Why it matters | Red flag |
|---|---|---|
| SKUs with gallery images vs SKUs with shelf crops | Identifies gallery-only SKUs that cannot be evaluated cross-domain | Large gap with no disclosure |
| Crops per SKU, by variant family | Shows near-duplicate coverage | Only one variant per family has crops |
| Packaging version table with dates | Prevents mixing old and new artwork under one ID | "Current images only" |
| Stores, fixtures and devices represented | Predicts robustness to deployment cameras | One store or one device |
| Label verification method | Indicates label noise rate | No audit sample or agreement figure |
| Rights basis for packshots and store photos | Determines whether training is permitted | Display-only syndication terms |
Ask for a datasheet alongside the data. A Data Card style summary of upstream sources, collection and annotation methods, intended use and known limitations makes these answers auditable [7].
Evaluating a candidate SKU recognition dataset
Evaluate purchased recognition data on a held-out set of your own shelf images before committing to volume. Build a small test set from the stores and cameras you actually deploy to, label it to the supplier's SKU IDs, and measure top-1 and top-5 retrieval accuracy against the supplier's gallery, broken out by variant family. Then report open-set behavior: how often crops of SKUs absent from the gallery are rejected instead of being matched to a lookalike.
Three diagnostics catch most problems early. First, compare accuracy on catalog-to-catalog retrieval with shelf-to-catalog retrieval; a large drop quantifies the domain gap you are paying to close. Second, run an embedding near-duplicate search across splits to find leaked crops. Third, sample the confusion matrix within GS1 bricks to confirm that errors cluster on true variants rather than on label mistakes.
For broader sizing questions, see how many images you need to train a computer vision model. If your problem is matching field photos of parts rather than packaged goods, visual part identification data is the closer guide, and text-to-SKU matching lives under industrial product matching data. The image data hub lists every image guide.
Source SKU recognition training data with SourceX
SourceX looks for US businesses that hold the data you describe, assesses the data and its licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Describe your SKU recognition data needs.
Frequently asked questions
Can I train SKU recognition with catalog images only?
You can bootstrap a gallery with catalog images only, but you cannot evaluate it honestly without labeled shelf crops. Heavy augmentation (perspective warps, blur, glare, color shifts) narrows the gap but does not reproduce occlusion by price rails or fridge-door reflections.
Do I need box annotations if I only care about recognition?
You need the box coordinates that produced each crop, even if your detector is trained elsewhere. They let you re-crop with your own padding and match the detector's output distribution at inference time.
Where does SourceX fit?
SourceX sources operational datasets from US companies on request; categories are not inventory and a request does not guarantee a match. It does not source scraped web content or generic CCTV or photos, and every release is approved by the supplying company.
Sources
- arXiv (Peng et al.), "RP2K: A Large-Scale Retail Product Dataset for Fine-Grained Image Classification" (2020). https://arxiv.org/pdf/2006.12634v1
- Mendeley Data, "Retail-YU: A Large-Scale Dual-Domain Dataset for Fine-Grained Retail Product Recognition". https://data.mendeley.com/datasets/mmcf24t9vv/1
- Ultralytics, "SKU-110K Dataset". https://docs.ultralytics.com/datasets/detect/sku-110k
- arXiv, "Semi-supervised Learning for Dense Object Detection in Retail Scenes" (2021). https://arxiv.org/pdf/2107.02114
- GS1 Netherlands, "GPC in a nutshell" (2024). https://www.gs1.nl/media/sfpiaxye/gpc-in-a-nutshell_jun24-def.pdf
- USENIX Security 2023, "Extracting Training Data from Diffusion Models" (2023). https://www.usenix.org/conference/usenixsecurity23/presentation/carlini
- arXiv (Google Research, FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.