Industry-specific operational data
Product categorization and taxonomy-mapping data for retail catalogs
Quick answer
Product categorization training data is a set of product records whose category labels were verified by people who own the outcome: merchandisers, catalog managers or marketplace onboarding teams. The useful version carries more than a final leaf. It records the supplier's original category, the rule or model suggestion, every human correction, and mappings to standard taxonomies such as GS1 GPC and Google's product taxonomy, each pinned to a dated taxonomy version so you can rebuild labels for the version you deploy against.
By SourceX Editorial · Updated
This guide is for teams that build categorization and mapping models for PIM, syndication and marketplace onboarding. It sits in the industry-specific operational data hub and stays narrowly on category assignment and cross-taxonomy mapping labels. Attribute values have their own guide on product attribute extraction data, and catalog licensing in general is covered on licensing product catalogs and descriptions.
Why public product datasets rarely cover taxonomy mapping
Public product datasets mostly give you one taxonomy and one label per item, which is not enough to train a mapping model. MAVE, for example, spans 2.2 million products across 1,257 categories built from marketplace product pages [1][2]. It is a strong attribute-extraction resource, but its categories reflect a single marketplace's tree at one point in time, and the rights in the underlying page content need a separate check from the repository license.
What mapping models need is the same product labeled in two or more trees at once. A syndication vendor sends a supplier item into a retailer's internal hierarchy, a GS1 GPC brick for GDSN, and a marketplace category. That co-labeled structure exists inside retailers and marketplaces as a byproduct of onboarding work, not in benchmark corpora.
There is also a provenance gap. A public label tells you the final answer, not who decided it or how often it was wrong first. Without that history you cannot separate easy items from the ambiguous ones where models actually fail.
The labels worth requesting, and where each comes from
The most useful labels are the ones that show each step of the decision, from supplier claim to merchandiser sign-off. Each step lives in a different system, so ask explicitly for each.
- Supplier-assigned category. The value in the vendor's item setup sheet, EDI 832 price/sales catalog, or GDSN item feed. It is often wrong or coarse, which makes it a realistic model input.
- Rule or model suggestion. Output of the PIM's auto-classification rules, keyword maps or an earlier model, with a confidence score where one was logged.
- Merchandiser correction. The change event: old value, new value, user role, timestamp and, if captured, a reason code such as "wrong department" or "too generic."
- Final internal category. The leaf node in the retailer's own hierarchy (department, class, subclass, or deeper), plus the hierarchy version.
- GS1 GPC brick. GPC organizes products as segment, family, class and brick; the brick is the working code used in GDSN item data, so request the brick code and the GPC release it came from.
- Marketplace or channel category. For Google Shopping feeds,
google_product_categoryholds a node from Google's predefined product taxonomy, whileproduct_typecarries the merchant's own category path. Other marketplaces keep their own browse or category trees with their own IDs.
The difference between product_type and google_product_category in a merchant feed is itself a mapping pair. A feed that carries both, with corrections to either, already holds the kind of internal-to-external alignment that mapping models learn from.
Why correction histories are the high-value signal
Correction histories are the most informative part of a categorization dataset because each wrong-to-right pair marks a boundary the first classifier could not see. A phone case first filed under "Mobile Phones" and moved to "Mobile Phone Accessories" teaches the exact confusion your model is likely to repeat.
Ask for corrections as an event log, not as an overwritten final column. In most PIM and MDM tools this lives in an audit or history table keyed by item ID, with field name, previous value, new value, actor and time. If the supplier's tool only keeps the latest value, a weekly snapshot diff is a usable fallback, though it loses intermediate states.
Watch for three failure modes in correction logs:
- Bulk reclassifications. A taxonomy release that moves 40 subclasses at once looks like thousands of corrections. Tag these as migration events so they do not dominate training.
- Rollback churn. An item moved and moved back within a day usually reflects a mistaken bulk edit, not a hard example.
- Unattributed changes. Edits by integration service accounts are not human verification. Keep the actor type field so you can filter them.
Versioning taxonomies so labels stay rebuildable
Every label is only valid against a specific taxonomy version, so a categorization dataset needs dated mapping tables, not just category strings. GS1 publishes periodic GPC releases, marketplace category trees change as channels add and retire nodes, and retailer hierarchies change with every line review. Ask each supplier which release every label was assigned against.
Do not assume a node's code encodes its parents; codes can survive a move to a new parent. Ship the full node table for each version: node ID, parent ID, label, level, valid-from and valid-to dates. Then add a crosswalk table that maps nodes between versions, including splits, merges and retirements.
With those tables you can re-project historical labels onto the taxonomy version your customer runs today. Without them, a model trained on 2024 labels quietly scores itself against categories that no longer exist.
Illustrative record and crosswalk schema
A deliverable that supports training and evaluation pairs an item-level record with a versioned taxonomy and crosswalk tables. JSON Lines works well for item records because each line is one UTF-8 JSON object, which keeps streaming and sharding simple [4]. If you would rather have co-labeled records like this located among US businesses that hold them than build the pipeline yourself, you can describe the records to SourceX.
Illustrative example: invented to show structure; it does not describe an available dataset.
{"item_id": "RTL-000418237",
"gtin": "redacted-or-hashed",
"title": "Silicone case for 6.1-inch phone, matte black",
"supplier_category": "Electronics > Phones",
"suggested_category": {"node_id": "H-2024B-4410", "source": "rules_v7", "score": 0.62},
"corrections": [
{"from": "H-2024B-4410", "to": "H-2024B-4472", "actor_type": "merchandiser",
"reason_code": "accessory_not_device", "ts": "2025-03-11T14:02:00Z"}
],
"final_internal": {"node_id": "H-2024B-4472", "taxonomy_version": "retailer-2024B"},
"gpc": {"brick": "BRICK_CODE", "gpc_release": "2024-11"},
"channel": {"google_product_category": "GOOGLE_ID", "product_type": "Phones > Cases"},
"verified": true}
Companion tables to request alongside the records:
| Table | Key fields | Why it matters |
|---|---|---|
| taxonomy_nodes | taxonomy, version, node_id, parent_id, label, level, valid_from, valid_to | Rebuilds the hierarchy for hierarchical loss and per-level metrics |
| version_crosswalk | taxonomy, from_version, from_node, to_version, to_node, change_type (split, merge, rename, retire) | Re-projects old labels onto the deployed version |
| cross_taxonomy_map | source_taxonomy, source_node, target_taxonomy, target_node, mapping_type (exact, broader, narrower), verified_by, verified_at | Trains and evaluates node-to-node mapping separately from item classification |
| correction_events | item_id, field, from, to, actor_type, reason_code, ts, batch_id | Separates hard examples from bulk migrations |
The mapping_type column matters because many internal subclasses map to a broader GPC brick or marketplace node. Treating every mapping as exact inflates accuracy on the mapping task.
Evaluating long-tail hierarchical classifiers
Evaluate categorization per hierarchy level, with macro-averaged metrics and a held-out category set, because a few head categories can hide failure across the long tail. Many leaf nodes hold only a handful of products, so micro-averaged accuracy mostly measures the head.
A practical evaluation protocol:
- Per-level scores. Report accuracy or F1 at department, class and leaf separately. A model that gets the department right but the leaf wrong is less costly than one that crosses departments.
- Macro averaging with minimum support. Macro-F1 over leaves weights rare categories equally; report how many leaves fall below a minimum example count.
- Held-out categories. Withhold whole leaves, not just items, to test whether the model can place items into nodes it has rarely seen, which mirrors what happens after a taxonomy release.
- Correction-derived hard set. Build a separate slice from items that a merchandiser corrected. This is where production errors live.
- Honest error bars. On small slices, normal-approximation confidence intervals tend to be too narrow; a recent position paper argues against CLT-based intervals with fewer than a few hundred data points [3]. Use bootstrap or exact intervals for small leaf slices.
Keep evaluation suppliers and training suppliers separate where you can. Labels from one retailer's merchandising team share conventions, and testing on the same team's work overstates transfer to a new customer's catalog.
Rights and privacy questions specific to catalog labels
Category labels and the product content they describe often carry different rights, so diligence has to split them. Manufacturer descriptions and images may be licensed to the retailer for display only, while the category assignments and correction logs are usually the retailer's own work product. Confirm both before assuming a dataset can be used for training.
Questions to put to any supplier:
- Who authored the titles, descriptions and images, and do the supplier agreements allow use beyond display?
- Are the category assignments and audit logs the supplier's own records, and can they be licensed for model training?
- Do item records include GTINs or supplier identifiers that reveal confidential vendor relationships or cost data?
- Do correction events carry employee names or user IDs that should be replaced with role labels?
- Is the internal taxonomy itself treated as confidential, and if so, may node labels be shared or only coded?
The product catalogs insight covers why catalog content is valuable for AI, and the data RFI guide shows how to collect answers like these from several suppliers before an RFP.
How this differs from matching, search and spend classification
Categorization assigns an item to a node; adjacent tasks need different labels and should be sourced separately. Product matching across sellers needs item-to-item identity pairs. Product search relevance data needs queries and graded judgments. Spend classification with UNSPSC and GS1 works from purchase lines and invoices rather than sell-side catalog records.
Returns data is another neighbor. Miscategorized items can drive "not as described" returns, and returns and RMA reason data can help prioritize which category boundaries to fix first.
Retailers that hold the full chain are usually multi-category sellers with an active merchandising team and channel syndication; see the buyer pages for e-commerce and multi-brand retail.
Sourcing verified categorization and mapping records
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery, and the supplying company approves every release; a request does not guarantee a match. Describe the taxonomies, versions and correction history you need on the SourceX buyers page.
Frequently asked questions
Can I train a mapping model from a single-taxonomy dataset?
You can train a classifier for that taxonomy, but node-to-node mapping needs items labeled in at least two trees, or a verified crosstaxonomymap. Without co-labels the model has to infer alignment from label text, which breaks where names match but scope differs.
How should I handle items that a merchandiser never reviewed?
Keep them, but flag them as unverified and exclude them from evaluation. Rule-assigned labels that nobody checked reflect your incumbent system's errors, so training on them unflagged teaches the model to reproduce those errors.
Is GPC enough as a universal pivot taxonomy?
GPC is a useful pivot for GDSN-connected categories, and its brick-level attributes help disambiguate. It is rarely as fine-grained as a retailer's merchandising leaves, so expect many-to-one mappings and record them as broader rather than exact.
Sources
- Google Research, "MAVE: A Product Dataset for Multi-source Attribute Value Extraction" (2022). https://research.google/pubs/mave-a-product-dataset-for-multi-source-attribute-value-extraction/
- arXiv, "MAVE: A Product Dataset for Multi-source Attribute Value Extraction (arXiv 2112.08663)" (2021). https://arxiv.org/abs/2112.08663v1
- ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
- jsonlines.org, "JSON Lines". https://jsonlines.org/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.