Tables, time series and transactional data
Product Matching Data Across Sellers and Catalogs
Quick answer
A useful product matching dataset is a set of offers from many sources, grouped into clusters that each represent one real product, with pairwise labels derived from those clusters. It needs reliable identifiers (GTIN, MPN plus brand, supplier-to-manufacturer cross-references) to seed labels, deliberate hard negatives such as size, color and pack variants, and test entities that never appear in training. Public benchmarks cover web retail well but rarely reflect B2B catalogs, missing identifiers or your own variant rules.
By SourceX Editorial · Updated
This page is for ML engineers building offer matching, catalog deduplication or price comparison. It sits in the tables, time series and transactional data hub and is distinct from entity resolution over people and company master data, which relies on names, addresses and registration numbers rather than product identifiers and attributes.
What a product matching record actually contains
The core unit is an offer, not a product: one seller's listing at one point in time, carrying a title, description, attribute table, price, identifiers and the source it came from. Matching models learn to decide whether two offers describe the same real-world product, so the dataset must preserve the raw, messy offer text alongside any normalized fields. Stripping offers down to clean attributes removes exactly the noise the model has to handle in production.
Two label formulations are common, and good datasets support both. Pairwise binary labels (match or non-match) train cross-encoders such as Ditto, while cluster labels (every offer assigned to a product ID) support multi-class and contrastive training. Some academic product benchmarks publish both formulations for the same offers, which lets teams compare approaches on identical data. If a supplier only gives you pairs, ask whether the underlying clusters exist, because pairs can always be generated from clusters but not reliably the reverse.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"offer_id": "src07-000418823",
"source_id": "src07",
"source_type": "distributor_catalog",
"captured_at": "2025-11-03",
"title": "Hex Cap Screw 3/8-16 x 1-1/2 Gr8 Zinc Yellow 50pk",
"brand_raw": "",
"mpn_raw": "HCS38-16112G8Y",
"gtin": null,
"attributes": {"thread": "3/8-16", "length_in": "1.5", "grade": "8", "finish": "zinc yellow", "pack_qty": "50"},
"price": {"amount": 18.40, "currency": "USD", "uom": "PK"},
"cluster_id": "P-0091273",
"label_source": "manufacturer_cross_reference",
"label_confidence": "verified_by_reviewer",
"variant_of": "P-0091270",
"variant_axis": "pack_qty"
}
The fields that matter most for training are cluster_id, label_source, and the variant fields. label_source tells you whether a match came from a shared GTIN, a cross-reference table, or a human judgment, and those three have very different error profiles. variant_of and variant_axis turn near-duplicates into explicit hard negatives instead of leaving them as silent label noise.
Identifiers: where labels come from and where they run out
GTINs are the cheapest source of positive labels, but they only work where they exist and where they were assigned correctly. Under GS1 allocation rules, each size, color and size-color combination generally gets its own GTIN, net content changes require a new GTIN, and each packaging level from consumer unit to case and pallet carries its own number [4]; check the current GS1 GTIN Management Standard for edge cases. That makes GTIN equality a strong match signal and GTIN inequality a strong variant signal, as long as sellers report the GTIN at the right packaging level.
In practice, identifier coverage is uneven, and the gaps are not random:
- B2B and industrial catalogs often carry only a manufacturer part number and a distributor SKU, with no GTIN. Matching then depends on MPN normalization (stripping hyphens, suffixes and packaging codes) plus brand, which is frequently blank or abbreviated.
- Marketplace offers sometimes reuse one GTIN across colors, or attach a case-level GTIN to a single-unit listing. These become false positives if GTIN equality is treated as ground truth.
- Private label and bundles have retailer-assigned codes that match nothing outside that retailer.
- Cross-reference tables (supplier part to manufacturer part, competitor part to own part) are a rich label source in distribution, but they often encode "acceptable substitute" rather than "identical product." Ask the data holder which relationship each row asserts. For this specific case, see industrial part cross-reference and description-to-SKU matching data.
A sound dataset records identifier coverage per source, so you can measure how your model performs on the offers that have no identifier at all. That slice is usually where a production matcher earns its keep.
Hard negatives: variants, packs and condition
Hard negatives are the non-matching pairs that look almost identical, and a dataset without enough of them trains a model that matches anything with a similar title. Benchmark authors at the University of Mannheim treat the share of these corner cases as an explicit dimension, because matching systems behave very differently as it changes [5]. Your request should specify which variant axes count as different products in your business, because that is a policy decision, not a data property.
Typical axes to label explicitly:
| Variant axis | Example pair | Usually a match? | Why it trips models |
|---|---|---|---|
| Color or finish | Same drill, black vs. red housing | No (own GTIN under GS1 rules) | Titles differ by one token |
| Size or dimension | 3/8-16 x 1-1/2 vs. 3/8-16 x 1-3/4 | No | Numbers tokenized poorly |
| Pack quantity | Single unit vs. 50-pack | No for inventory; maybe yes for price per unit | Same MPN stem with pack suffix |
| Net content | 500 ml vs. 750 ml | No (new GTIN under GS1 rules) | Units in free text |
| Condition | New vs. refurbished vs. open box | Policy-dependent | Often only in a condition field |
| Model year or revision | Rev A vs. Rev B board | Usually no | Revision hidden in description |
| Region or voltage | US plug vs. EU plug | No | Same brand, model name and photo |
If you train for price comparison, you may want a three-way label (exact match, same product different pack, different product) rather than binary. Ask for the variant relationship as data, not as a collapsed binary label, so you can choose the policy later.
Where public product matching benchmarks fall short
Public benchmarks are useful for method selection but are a weak proxy for a production catalog. A 2024 critical re-evaluation of learning-based matching benchmarks, many of them product-based, found quality issues that limit what high scores actually demonstrate [2]. Widely reused pairs such as Abt-Buy and Amazon-Google are small and long studied, so leaderboard progress on them says little about your catalog.
Newer web-derived benchmarks from the Web Data Commons project build offers from schema.org product markup across many e-shops and control for corner cases, unseen entities and training-set size. They are still bounded snapshots of web retail listings rather than distributor feeds, ERP item masters or procurement catalogs. Weak performance on unseen entities, a recurring finding in that line of work, is exactly the condition a live catalog creates every day as new products arrive.
The gaps a buyer usually needs to fill:
- Domain: industrial, healthcare supply, automotive aftermarket or grocery B2B, where titles are code-heavy and GTINs are sparse.
- Unseen-entity evaluation: a test set whose product clusters are entirely absent from training.
- Temporal drift: offers captured over several periods, so titles, prices and discontinued items reflect real churn.
- Your variant policy: labels that follow your definition of "same product," not a benchmark author's.
How many labeled pairs you need, and how to spend them
Label volume matters less than label composition, but there are useful reference points. WDC Block publishes training sets of roughly 1K, 5K and 20K pairs specifically to measure how blocking quality changes with labeled data [1]. Run the same kind of size sweep on your own sample before buying at scale, because architecture choice (cross-encoder versus contrastive bi-encoder) changes how many labels you need.
Spend labels in this order: hard negatives from your variant axes first, then positives from offers without shared identifiers, then easy pairs last. Blocking and matching need different data, so separate them. A blocker needs high recall across the full candidate space, while a matcher needs dense coverage of the borderline cases the blocker lets through.
Specifying and accepting a product matching dataset
A product matching request should state the label unit, the identifier rules and the evaluation split before any price discussion. The checklist below turns that into acceptance criteria you can test on a sample.
Illustrative example: invented to show structure; it does not describe an available dataset.
Product matching data request checklist
- Sources: number of distinct sellers or catalogs, source types (retailer, marketplace, distributor, manufacturer), capture dates.
- Unit: raw offer records with source ID and capture date; cluster IDs and generated pairs, or both.
- Identifier coverage: percent of offers with GTIN, MPN plus brand, and neither, reported per source.
- Label provenance: share of labels from GTIN equality, cross-reference tables and human review; reviewer agreement on a double-labeled subset.
- Variant policy: axes treated as distinct products, with
variant_oflinks preserved. - Hard-negative share in train and test.
- Unseen-entity test split: product clusters absent from training.
- Leakage controls: no offer, and no near-duplicate offer text, appearing in both train and test.
- Format and delivery: Parquet or CSV with a data dictionary; see dataset delivery formats and schemas.
- Rights: confirmation the data holder can license offer text, attributes and prices for model training.
For quality acceptance, ISO/IEC 5259-3 sets requirements for managing the quality of data used in ML and is a reasonable frame for documenting label checks [3]. Our training data quality assessment guide covers sampling, contamination and coverage testing in more depth. For the table-level terms (grain, keys, history) to put in the license, see what to specify when licensing tabular data.
Adjacent tasks that share the same catalog data
Product matching often shares source data with three neighboring tasks, and it is cheaper to specify them together. Product attribute extraction and normalization data provides the structured attributes that make variant labels possible. Product categorization and taxonomy-mapping data helps blockers by narrowing candidates to the same category. Product search relevance data uses query judgments rather than offer pairs, but it draws on the same catalogs. If you need the catalogs themselves, the product catalogs and descriptions page describes that request.
Record each licensed source and its allowed uses in a training data use register, because offer data often comes from several holders with different terms.
How SourceX fits a product matching request
SourceX sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases. Relevant holders here are businesses with their own catalogs, item masters and cross-reference records, not scraped web listings: SourceX does not source scraped web content. Data is sourced on request rather than held in stock, so you describe the offers, identifiers and labels you need on the buyer request page, and a request does not guarantee a match.
Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery. Every release is approved by the supplying company, and nothing is contracted until a supplier agrees. Delivery runs through private, access-controlled workflows only after an executed agreement.
Request product matching data from US catalog holders
If your matcher needs offers, cross-references and variant labels from real business catalogs, describe the data rather than the companies that hold it. SourceX looks for US businesses that hold it, reviews rights, and manages licensing for AI teams wherever they are based. Describe the product matching data you need.
Frequently asked questions
Can I build product matching labels from GTINs alone?
Only partly. GTIN equality gives strong positives where GTINs are present and correctly assigned at the consumer-unit level, but B2B catalogs often lack them, and marketplace listings can misuse them. You still need human-reviewed labels for offers without identifiers and for variant pairs.
Should the test set include products seen in training?
Include both, and report them separately. Performance on unseen entities is a known weak point of current matchers, and it is the condition your catalog faces as new products arrive.
Are offer prices useful features or a leakage risk?
Both. Price helps separate pack-size variants, but if the dataset groups offers by price band during labeling, the model can learn the labeling shortcut. Ask how clusters were formed before using price as a feature.
Sources
- University of Mannheim, Data and Web Science Group, "WDC Block: A large Blocking Benchmark released". https://www.uni-mannheim.de/dws/news/wdc-block-a-large-blocking-benchmark-released/
- ICDE 2024 (Papadakis et al.), "A Critical Re-evaluation of Benchmark Datasets for (Deep) Learning-Based Matching Algorithms" (2024). https://helios2.mi.parisdescartes.fr/~themisp/publications/icde24-dlmatching.pdf
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-3:2024 Data quality for analytics and ML, Part 3: Data quality management requirements and guidelines" (2024). https://www.iso.org/standard/81092.html
- GS1, "GS1 GTIN Management Standard". https://www.gs1.org/standards/gs1-gtin-management-standard
- University of Mannheim (Peeters et al.), "WDC Products: A Multi-Dimensional Entity Matching Benchmark" (2024). https://webdatacommons.org/largescaleproductcorpus/wdc-products/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.