Skip to content

Tables, time series and transactional data

Product Attribute Extraction Data: Sourcing Verified Attribute-Value Labels from PIM and ERP Records

Quick answer

The strongest product attribute extraction dataset pairs raw supplier text (titles, descriptions, spec sheets) with attribute values that a business verified in its PIM or ERP, plus the taxonomy and unit vocabulary used to normalize them. Public benchmarks such as MAVE are useful for comparison but come from scraped marketplace pages. For commercial training, license operational product records with documented label provenance, B2B spec fields, and clear rights in the underlying descriptions.

By SourceX Editorial · Updated

What a product attribute extraction dataset needs to contain

A usable dataset links three layers for each SKU: the unstructured source text, the verified attribute-value pairs, and the schema that defines which attributes apply to which category. Without the third layer you cannot tell a missing value from an attribute that does not apply, which is the most common silent label error in this task. Extraction models learn span or generation targets from layer two, while normalization models learn the mapping from surface strings ("1/2 in", "12.7mm", "half inch") to canonical values.

In operational systems these layers live in different places. The PIM (Akeneo, Salsify, inRiver, Stibo STEP, Syndigo and similar) holds the attribute families, option lists and channel-ready values. The ERP item master (SAP MARA/MARM, Oracle EBS MTL_SYSTEM_ITEMS, NetSuite item records) holds units of measure, weights, manufacturer part numbers and purchasing descriptions. Supplier onboarding files, often spreadsheets or BMEcat/ETIM exports in industrial distribution, hold the raw text that the enrichment team cleaned.

Ask suppliers which system was the system of record for each attribute and whether values were set by a person, imported from a manufacturer feed, or filled by an earlier model. Model-filled values recycled as labels will teach your model the previous model's errors.

Why public benchmarks like MAVE are a reference, not a training source

MAVE is a widely used public reference: Google Research released 2.2 million products with 3 million attribute-value annotations across 1,257 categories, built from Amazon product pages and published at WSDM 2022 [1]. The authors describe it as the largest attribute value extraction dataset by attribute-value examples and include a zero-shot test set for unseen attributes [2]. That makes it valuable for architecture comparisons and for reporting results reviewers recognize.

The limits for a commercial buyer are structural. MAVE was built from marketplace product pages [2], and a dataset built that way does not necessarily convey rights in the product descriptions it contains, so "MAVE dataset commercial use" questions turn on the underlying page content, not only on the dataset's own license. Broad audits show this is a general problem: the Data Provenance Initiative found license omission above 70% and license error rates above 50% for popular datasets on hosting sites [6].

There is also a distribution gap. Consumer marketplace listings over-represent apparel, electronics and home goods, and their attributes skew toward color, size and material. A distributor enriching fasteners, valves or electrical components needs thread pitch, pressure rating, IP rating and material grade, which public consumer benchmarks rarely cover.

Verified PIM and ERP attributes versus scraped labels

Attributes confirmed in a PIM workflow generally make cleaner labels than values inferred from scraped text, because a merchandiser or data steward accepted them against a defined option list. That is a working hypothesis to test, not a guarantee: PIM data has its own failure modes. Check for these before you treat any field as ground truth.

  • Default-value pollution: fields set to a family default ("Color: Black") when nobody checked.
  • Stale values: a product reformulated or re-specified while the PIM kept the old value; compare against effective dates in the ERP.
  • Free-text leakage: option-list attributes overridden with free text in one channel export.
  • Inherited values: variant-level attributes copied from the parent model when they differ by SKU.
  • Label-text leakage: descriptions generated from the attributes themselves, which makes extraction trivially easy and inflates offline scores.

The last point matters most. If a channel template wrote "Stainless steel 304, 1/4-20 UNC, 2 in length" into the title from structured fields, your model is learning to copy a template. Ask for the provenance of each text field (manufacturer feed, supplier spreadsheet, copywriter, template) and hold out records where text was written independently of the attribute values. For time-stamped values, apply the as-of logic described in our guide to point-in-time correct training data.

B2B product specification data: units, tolerances and part numbers

B2B product specification data is harder than consumer catalog data because values are numeric, unit-bearing, toleranced and tied to external standards. A single attribute such as "operating temperature" may appear as "-40 to 85 C", "-40°F~185°F" or "Industrial temp range" in different supplier files. Your labels need the numeric range, the unit and the qualifier separately.

Fields worth requesting explicitly:

  • Manufacturer name and manufacturer part number (MPN), plus the distributor's internal item number and GTIN where assigned.
  • Unit of measure for each numeric attribute, with the code list used (UN/CEFACT Recommendation 20 codes, UCUM, or an internal list).
  • Tolerances and ranges stored as structured min, max and nominal values, not strings.
  • Standards references (ASTM, DIN, ISO, UL, NEMA) as separate fields from the free-text description.
  • Classification codes: GS1 GPC brick, UNSPSC, ETIM or ECLASS class, depending on the vertical.

Master data quality conventions help here. ISO 8000 defines requirements for master data quality, including ISO 8000-115 on exchanging quality identifiers with syntactic, semantic and resolution requirements [5]. If a supplier already follows ISO 8000 or ECLASS-style property definitions, mapping their attributes to your target schema is much cheaper.

Normalization targets: canonical units, value vocabularies and taxonomies

Normalization training data is the mapping from raw values to a controlled target, so the target vocabulary must ship with the records. That means the category taxonomy, the per-category attribute list, allowed values for enumerated attributes, and canonical units with conversion rules.

Common taxonomies include GS1 GPC, which organizes products into a segment, family, class and brick hierarchy with attributes defined at brick level [4], and Google's product taxonomy, where the google_product_category attribute must use a predefined category by numeric ID or full path while product_type carries the merchant's own labels [3]. Pairs of product_type and google_product_category in a merchant's feed history are themselves useful taxonomy-alignment labels. For classification models trained on these codes, see product and spend classification training data.

Request the taxonomy version and change log. A brick or category retired mid-history will appear as label noise unless you can map old codes forward.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "item-000418",
  "source_text": {
    "title": "SS HEX CAP SCREW 1/4-20 X 2 GR 18-8",
    "description_origin": "supplier_spreadsheet",
    "description": "Hex head cap screw, 18-8 stainless, full thread, 1/4in-20 UNC, 2in long"
  },
  "category": {"taxonomy": "GS1 GPC", "version": "2025-11", "code": "10005779"},
  "attributes": [
    {"name": "thread_size", "raw": "1/4-20", "value": "1/4-20 UNC", "source": "pim_steward", "verified_at": "2025-03-12"},
    {"name": "length", "raw": "2in", "value": 50.8, "unit": "MMT", "source": "erp_item_master"},
    {"name": "material", "raw": "18-8 stainless", "value": "Stainless steel 18-8", "source": "pim_steward"},
    {"name": "head_style", "raw": "HEX", "value": "Hex", "source": "pim_steward"}
  ],
  "not_applicable": ["voltage", "ip_rating"],
  "mpn": "redacted-in-sample",
  "text_generated_from_attributes": false
}

The not_applicable list and the text_generated_from_attributes flag are the two fields most often missing from catalog exports and the two that most change evaluation results.

Rights in product descriptions and manufacturer content

The company that holds product records does not always own every piece of text in them. Manufacturer-supplied descriptions, marketing copy from content syndication networks, and images may carry the manufacturer's own copyright or be distributed under syndication terms that restrict reuse. Attribute values themselves (a thread size, a voltage) are largely factual, but the descriptive prose you need as model input is where rights questions concentrate.

Ask the supplying company to identify, per text field, whether it authored the content, received it under a manufacturer or syndication agreement, or copied it from a public page. Exclude or separately clear fields it cannot account for. If you train a general-purpose model placed on the EU market, Article 53(1)(c) of the AI Act requires a copyright compliance policy, including respect for text-and-data-mining reservations under Article 4(3) of Directive (EU) 2019/790 [8]; as of October 2026 those duties apply, with AI Office enforcement for new models from August 2026.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Buyer checklist for product attribute data

Use this checklist when scoping a request or reviewing a sample. Documentation in the style of Data Cards, covering upstream sources, collection and annotation methods, intended use and known limits, should answer most of these items before you see records [7].

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckWhat to ask forWhy it matters
Label provenanceSource system and actor per attribute (steward, feed, model)Model-filled labels propagate errors
Applicability schemaCategory-to-attribute map with required and optional flagsSeparates "missing" from "not applicable"
Text independenceFlag for template-generated titles and descriptionsPrevents leakage and inflated scores
Units and vocabulariesUoM code list, enumerations, conversion rulesDefines normalization targets
Taxonomy versionVersion IDs and code change logAvoids drift-induced label noise
B2B fieldsMPN, GTIN, tolerances, standards referencesCovers industrial extraction cases
Rights per fieldAuthorship or license basis for each text fieldManufacturer copy may not be the holder's to license
Personal dataConfirmation that supplier contacts and buyer notes are removedCatalog exports often carry them
Holdout designCategories or attributes reserved for zero-shot testingMirrors MAVE-style unseen-attribute evaluation [2]

For matching the same product across sellers, which uses overlapping fields but a different label type, see product matching data across sellers and catalogs. For image-based attribute labels, see product catalog photo datasets.

How SourceX sources product attribute records

SourceX sources operational datasets from US companies on request; a request can describe product and item-master records of the kind distributors, manufacturers and retailers keep in PIM and ERP systems. You describe the data you need, such as categories, attributes, label provenance and text fields, not the businesses that hold it; SourceX looks for US companies holding that data, and every release is approved by the supplying company. Nothing is held in stock, and a request does not guarantee a match. SourceX does not source scraped web content.

Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as supplier contact names, emails and phone numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Broader context on catalog licensing sits on license product catalogs and descriptions and in what makes product catalogs valuable for AI. You can start a request on the SourceX buyer page, or return to the structured data buyer's guide and the AI data hub.

Request verified product attribute data for extraction models

SourceX sources operational datasets from US companies on request, rights-reviews each dataset, and delivers it under a license that defines records, uses, term and delivery, after the supplying company approves the release. Describe the categories, attributes and text fields your extraction or normalization model needs on the SourceX buyer page.

Frequently asked questions

Can I use MAVE to train a commercial attribute extraction model?

Check the dataset's own terms and, separately, the rights in the underlying product pages it was built from [2]. A common pattern is to use MAVE for benchmarking and architecture selection, then train production models on licensed records whose text provenance is documented.

How many labeled products do I need per category?

It depends on how many attributes the category has and how skewed their values are. A practical approach is to request a sample covering your long-tail categories first, measure per-attribute recall on a held-out slice, and size the full request from the attributes that underperform.

Is ERP item-master data enough on its own?

Usually not. ERP records carry units, weights, MPNs and short purchasing descriptions, but rich attribute values and option lists usually live in the PIM. The best training sets join both on item number, with the join keys documented.

Sources

  1. Google Research, "MAVE: A Product Dataset for Multi-source Attribute Value Extraction" (2022). https://research.google/pubs/mave-a-product-dataset-for-multi-source-attribute-value-extraction/
  2. arXiv, "MAVE: A Product Dataset for Multi-source Attribute Value Extraction (arXiv:2112.08663v1)" (2021). https://arxiv.org/abs/2112.08663v1
  3. Google Merchant Center Help, "Google product category [google_product_category]". https://support.google.com/merchants/answer/6324436?hl=en-GB
  4. GS1 Sweden, "Global Product Classification (GPC)". https://gs1.se/en/standards-and-services/global-product-classification-gpc/
  5. Standards Council of Canada, "ISO 8000-115:2024 Data quality - Master data: Exchange of quality identifiers" (2024). https://scc-ccn.ca/standardsdb/standards/8188648
  6. arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  7. arXiv (Pushkarna, Zaldivar, Kjartansson, Google Research), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  8. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data