Skip to content

Image data

Product Catalog Photo Datasets for E-commerce Vision and Attribute Extraction

Quick answer

A usable product image dataset for AI training is more than a folder of packshots. It joins each photo to a stable product ID, a versioned category taxonomy and verified attributes such as color, material and pattern. It also carries rights that cover training: photographer copyright, model releases for on-model shots, and a license naming the allowed uses. Many public sets are fashion-specific or carry restrictive terms, so teams building visual search, attribute extraction or VLM catalog enrichment usually license photos from the businesses that commissioned them.

By SourceX Editorial · Updated

What catalog photos train, and which shot types matter

Catalog photos train four model families: categorization, visual search, image generation and image-to-attribute extraction. Each one needs a different mix of shots. Researchers already use curated online product images per SKU as the clean reference domain for recognition and retrieval [1]. That is the same role a packshot plays when you match a user photo or shelf crop against a catalog. For the shelf side of that pairing, see SKU recognition with shelf crops and catalog reference images.

  • Packshots (white or neutral background, centered, consistent lighting) work best as retrieval anchors and as clean inputs for category classifiers.
  • Multi-angle sets (front, back, side, detail, label close-up) give attribute models the views they need for fields like closure type or care-label material.
  • Lifestyle and on-model images add context, clutter and pose variation, which helps when inference images come from users rather than studios.
  • Variant images (one shot per color or finish) show the model which attributes vary within a product family.

A catalog that only has packshots trains a model that does well in the studio and poorly on user uploads. Ask for the shot-type distribution per category before you commit.

The image-to-attribute label spec

Attribute extraction from images is a multi-label problem: one image carries several attribute values at once, and the label spec has to say so. The iMaterialist Fashion Attribute dataset shows the pattern at scale. It has more than one million expert-labeled images and 228 fine-grained attributes organized into 8 groups [3]. Smaller sets packaged for VLM fine-tuning follow the same idea. One public set pairs 34,493 garment photos with nine attributes, but as of October 2026 it is gated and lists its license as "other" [4], so check the terms before you use it commercially.

Text-side benchmarks such as MAVE (2.2 million products, 3 million attribute-value annotations, 1,257 categories) show how attribute values get grounded in product-page sources [5]. They are text, not photos. The structured counterpart of this page is product attribute extraction and normalization data. If your target is JSON output from a VLM, pair the images with structured-output fine-tuning examples.

Taxonomy is where most catalog datasets break. Google's google_product_category accepts only values from Google's predefined list, given as a numeric ID or a full path. Merchant-defined labels go in product_type [6]. Record which taxonomy each record uses and which version. Categories get renamed and merged over time, and a model trained on one version will silently mislabel against the next.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "image_id": "img_000418_v2_side",
  "product_id": "sku_7731-BLK",
  "parent_product_id": "sku_7731",
  "shot_type": "multi_angle_side",
  "background": "white_seamless",
  "on_model": false,
  "resolution_px": [2400, 3000],
  "format": "JPEG",
  "taxonomy": { "system": "google_product_category", "version": "2021-09-21", "id": 187 },
  "merchant_product_type": "Footwear > Boots > Chelsea",
  "attributes": {
    "color": ["black"],
    "material_upper": ["leather"],
    "closure": ["elastic_gusset"],
    "toe_shape": ["round"]
  },
  "attribute_source": "merchant_entered",
  "verified_subset": false,
  "rights": { "copyright_holder": "commissioning_brand", "model_release": "n/a", "ai_training_permitted": true }
}

Merchant-entered attributes are noisy, so verify a subset

Merchant feeds are written to pass listing validation, not to serve as ground truth. Expect the usual failures: missing values, free-text colors ("midnight", "ink"), a material copied across every variant, and wrong categories chosen to win placement. Training on that noise is often acceptable. Evaluating on it is not.

Ask for an expert-verified subset, stratified by category and attribute, that you keep as a held-out test set. Each record should carry an attribute_source field (merchant-entered, model-predicted, human-verified), so you can weight or filter labels. If no verified subset exists, budget for one through expert annotations and labels, and see pre-labeled vs raw image datasets for the tradeoff.

Photo rights differ from catalog-text rights

Product photos carry rights that listing text does not. The photographer or agency may own the copyright, people in on-model shots have publicity and release terms, and logos on the products are third-party trademarks. A retailer that licenses you its catalog may not hold training rights in images an outside studio shot under a contract that never mentioned AI. Ask for the work-for-hire or assignment terms, and confirm that model releases cover machine-learning use, not just advertising.

Scraping marketplace images does not solve this. Public image collections often mix licenses [2], and an audit of 1,800+ text datasets found licenses on popular hosting sites omitted more than 70% of the time and wrong more than 50% of the time [7]. For generative use there is a second risk: diffusion models can regenerate individual training images, including trademarked logos [8]. Your license should say whether generation is an allowed use. For public options with clearer terms, compare image and caption datasets you can train on commercially.

Buyer checklist for a product photo request

Describe the data, its links and its rights in one request, and ask suppliers for documentation in a data-card style (sources, collection, annotation, intended use) [9].

Illustrative example: invented to show structure; it does not describe an available dataset.

ItemWhat to specifyWhy it matters
Categories and taxonomyTaxonomy system, version, mapping table if customPrevents label drift across versions
Shot typesPackshot, multi-angle, lifestyle, on-model mix per categoryMatches the inference domain
Linkageimage_id to product_id to parent_product_id, variant keysEnables retrieval and variant-aware training
AttributesField list, allowed values, multi-label rules, attribute_sourceDefines the target for extraction
Verified subsetSize per category, who verified, agreement rateGives a trustworthy evaluation set
FilesFormat, minimum resolution, color profile, EXIF policySee resolution and compression requirements and EXIF metadata handling
RightsCopyright holder, work-for-hire terms, model releases, trademark notesConfirms training and generation rights
Allowed usesClassification, retrieval, generation, VLM fine-tuningGeneration needs explicit permission

Fashion buyers with garment-specific taxonomies and on-model issues should also read fashion and apparel image datasets. Listing text, specs and copy are covered by product catalogs and descriptions, and other photo types by images and inspection photos.

How SourceX handles product photo requests

SourceX sources operational datasets from US companies on request. You describe the photos, links and attributes you need, and SourceX looks for US businesses that hold them; categories are not inventory and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, every release is approved by the supplying company, and delivery happens under a license that defines records, uses, term and delivery. SourceX does not source scraped web content or generic photos, and it does not train models. Teams sourcing e-commerce training data can submit a buyer request. More image categories are on the image data hub and the AI data hub.

Request licensed product catalog photos

Describe the product categories, shot types, attributes and allowed uses you need. SourceX manages the find, assess, agree, transact and manage process with US suppliers, and nothing is contracted until a supplier agrees. Start a buyer request at SourceX.

Sources

  1. Mendeley Data, "Retail-YU: A Large-Scale Dual-Domain Dataset for Fine-Grained Retail Product Recognition" (2026). https://data.mendeley.com/datasets/mmcf24t9vv/1
  2. arXiv, "Can I use this publicly available dataset to build commercial AI software? -- A Case Study on Publicly Available Image Datasets" (2022). https://arxiv.org/abs/2111.02374v4
  3. ICCV Workshops 2019 (CVF Open Access), "The iMaterialist Fashion Attribute Dataset" (2019). https://openaccess.thecvf.com/content_ICCVW_2019/html/CVFAD/Guo_The_iMaterialist_Fashion_Attribute_Dataset_ICCVW_2019_paper.html
  4. Hugging Face, "Denali-AI/train-35k". https://huggingface.co/datasets/Denali-AI/train-35k
  5. Google Research, "MAVE: A Product Dataset for Multi-source Attribute Value Extraction" (2022). https://research.google/pubs/pub50791
  6. Google Merchant Center Help, "Google product category [google_product_category]". https://support.google.com/merchants/answer/6324436?hl=en-GB
  7. Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  8. USENIX Security 2023 (Carlini et al.), "Extracting Training Data from Diffusion Models" (2023). https://www.usenix.org/conference/usenixsecurity23/presentation/carlini
  9. Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data