Skip to content

Image data

Fashion and Apparel Image Datasets with Fine-Grained Attribute Labels

Quick answer

A usable fashion image dataset pairs each garment image with a multi-label attribute record (category, neckline, sleeve length, pattern, material, fit, closure) under a fixed taxonomy, separates flat-lay, ghost-mannequin and on-model views, and comes with rights that permit commercial training. Public sets such as iMaterialist prove the taxonomy depth is achievable, but their terms and partial releases rarely cover production use, so commercial teams usually license retailer or brand photography with model releases that name AI training.

By SourceX Editorial · Updated

What public fashion attribute datasets cover, and where they stop

Public fashion benchmarks are good for prototyping and pre-training experiments, but most are research artifacts, not commercial training sources. iMaterialist (iFashion-Attribute) labels over one million images with 228 fine-grained attributes in 8 groups, and each image can carry several labels at once [1]. Its authors report that pre-training on this fashion attribute set transferred better to fashion tasks than ImageNet pre-training [2], which is a strong argument for domain data rather than generic photo corpora.

The gaps show up when you move to production. Retrieval research still leans on FashionAI and DeepFashion attribute benchmarks, and FashionAI's full version was not publicly released [3]. Newer hub-hosted garment sets with brand, size and visible-defect labels can be gated behind contact sharing and published under non-standard licenses [4]. A broader audit of dataset hosting sites reported license omission rates above 70% and license error rates above 50% [6], so a "free" fashion set often has no defensible commercial-use chain.

For a wider view of commercially trainable image-text collections, see image and caption datasets you can train on commercially.

Attribute taxonomy: define it before you request a single image

The taxonomy is the product, because attribute labels are only comparable when every annotator and supplier uses the same definitions. Decide up front whether attributes are single-select (sleeve length: sleeveless, cap, short, elbow, three-quarter, long) or multi-label (pattern can be floral and striped on a color-blocked garment). Write a visual definition and a negative example for each value: "boat neck" versus "scoop neck" is where inter-annotator agreement collapses.

Plan for three structural issues that generic image specs miss:

  • Garment scope per image. On-model and street shots contain several garments. Require a bounding box or segmentation mask per garment, with attributes attached to the garment, not the image.
  • Unknown versus absent. A cropped photo may hide the hem. Use an explicit "not visible" value so the model is not trained to predict "no hem detail".
  • Taxonomy versioning. Merchandising teams rename and split categories each season. Store a taxonomy_version on every record and keep a mapping table between versions.

Label noise matters even in curated sets; benchmark test sets in widely used datasets carry an estimated average label error rate of at least 3.3% [7]. Budget a re-annotation audit on a held-out slice. The trade-offs between buying labeled data and annotating raw photography in-house are covered in pre-labeled vs raw image datasets.

Views, capture conditions and catalog metadata

Separate view types are worth more than a larger undifferentiated pile, because flat-lay, ghost-mannequin and on-model images teach different things. Flat-lay and ghost-mannequin shots isolate construction details (collar shape, pocket placement, stitching) on clean backgrounds. On-model images teach drape, fit and occlusion, which is what a visual search query from a customer phone photo looks like. Ask suppliers to label view type per image rather than inferring it later.

Retailer catalog photography usually carries structured metadata that doubles as weak labels: product title, merchandising category, color name, material composition from the care label, and size range. Those fields are noisy (marketing color names such as "oat" or "storm") but cheap to normalize against your taxonomy. The guide to product catalog photo datasets covers SKU-level pairing in more depth, and camera, lens and lighting diversity explains why studio-only data underperforms on user-generated queries.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueNotes
image_idimg_000412Stable, non-derivable from SKU or filename
view_typeon_modelflat_lay, ghost_mannequin, on_model, detail
garment_bbox[112, 64, 498, 902]One record per garment in multi-garment shots
categorydressFrom taxonomy_version 2026.2
necklinev_neckSingle-select; not_visible allowed
sleeve_lengththree_quarterSingle-select
pattern["floral"]Multi-label
material_declared100% viscoseFrom catalog care label, not visual
model_release_refrel_0193Present for every on-model image
release_ai_trainingtrueRelease text explicitly covers AI/ML training
logo_visiblefalseFlags third-party marks and prints
license_idlic_ALinks to allowed uses and term

Rights for on-model fashion photos

On-model fashion photos need two separate permissions: the photo's copyright license and a model release that covers AI and machine-learning training. Many legacy releases were written for print and e-commerce advertising and say nothing about model training; some current stock releases now address AI/ML training explicitly [5]. Ask for the release template, not a yes-or-no assurance, and confirm it covers the territory and the use you plan.

Printed graphics, brand logos and licensed characters on garments add a third layer. Diffusion models have been shown to memorize and regenerate individual training images, including photos of individual people and trademarked logos [8], so generative teams in particular should flag logo-bearing and print-heavy images and decide how to treat them. Faces are a related issue: see licensing image data that contains faces and the cluster guide to model and property releases for AI training images.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Buyer request checklist for fashion image data

A precise request gets better supplier answers than a category name. Use this checklist when describing the data to any supplier or broker.

Illustrative example: invented to show structure; it does not describe an available dataset.

  1. Task. Attribute recognition, visual search, outfit compatibility, or VLM captioning; each implies different labels.
  2. Taxonomy. Attach your attribute list, value definitions and not_visible rules; state single-select versus multi-label per attribute.
  3. Views. Minimum share of on-model, flat-lay, ghost-mannequin and detail images, labeled per image.
  4. Coverage. Categories, seasons, size ranges and body-type diversity on models; note gaps you will not accept.
  5. Metadata. Which catalog fields you want (title, color, material, category) and whether EXIF should be stripped; see EXIF metadata in image training data.
  6. Rights. Photo ownership, model release text that names AI training, treatment of third-party logos and prints.
  7. Quality. Agreement target on a dual-annotated slice and a re-annotation budget.
  8. Delivery. Image format and resolution floor, annotation format (COCO JSON or CSV), and a held-out evaluation split.

How SourceX helps fashion AI teams source image data

SourceX sources operational datasets from US companies on request, so a fashion image request is a search, not a catalog lookup, and it does not guarantee a match. Buyers describe the data they need, and SourceX looks for US businesses that hold it; every release is approved by the supplying company. The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery.

Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. SourceX does not source scraped web content or generic photos, and it does not train models. Teams buying adjacent data can start from licensed images and inspection photos, product catalogs and descriptions or the e-commerce buyer overview, and the image data hub and AI data hub cover related topics. To describe a specific taxonomy and view mix, submit a buyer request.

Request licensed fashion and apparel image data

Describe your garment categories, attribute taxonomy, view types and intended uses, and SourceX will look for US companies that hold matching photography. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, after supplier approval. Start a fashion image data request.

Sources

  1. CVF Open Access (ICCV Workshops 2019), "The iMaterialist Fashion Attribute Dataset" (2019). https://openaccess.thecvf.com/content_ICCVW_2019/html/CVFAD/Guo_The_iMaterialist_Fashion_Attribute_Dataset_ICCVW_2019_paper.html
  2. arXiv, "The iMaterialist Fashion Attribute Dataset (arXiv:1906.05750)" (2019). https://arxiv.org/pdf/1906.05750
  3. arXiv, "Attribute-Guided Multi-Level Attention Network for Fine-Grained Fashion Retrieval" (2023). https://arxiv.org/pdf/2301.13014
  4. Hugging Face, "Denali-AI/train-35k". https://huggingface.co/datasets/Denali-AI/train-35k
  5. pocstock, "Model Release". https://pocstock.com/legal/model-release
  6. arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  7. arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  8. USENIX Security 2023 (Carlini et al.), "Extracting Training Data from Diffusion Models" (2023). https://www.usenix.org/conference/usenixsecurity23/presentation/carlini

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data