Skip to content

Tables, time series and transactional data

Product and Spend Classification Training Data (UNSPSC, GS1 and Custom Taxonomies)

Quick answer

Spend classification training data is a set of real purchase-order, invoice and catalog line items, each paired with a taxonomy code (UNSPSC commodity, GS1 GPC brick, or a company's own category tree) plus a record of who assigned that code and when. The useful datasets keep the terse original line text, the side signals that classifiers depend on (supplier, GL account, material group, unit of measure), the taxonomy version, and a label-provenance flag per row. Public product data rarely provides any of that for procurement lines.

By SourceX Editorial · Updated

What a spend classification dataset must contain

A usable dataset contains the original line text, its structural context, a code at a stated taxonomy level and version, and provenance for that code. Line descriptions in AP and PO systems are short and abbreviation-dense ("SCR HEX M8X25 SS A2 BX100", "SVC FEE Q3 MAINT"), so the description alone often underdetermines the category. Practitioners rely on side signals: supplier identity, general ledger account, cost center, the ERP material group (in SAP, MATKL, which SAP itself uses as a requisition release criterion alongside plant and value [7]), unit of measure, unit price band and whether the line references a catalog item.

Treat these fields as a minimum when you write a request:

  • Line text: the raw short description and any long text, unnormalized, plus the source system and document type (PO, invoice, P-card, expense, catalog).
  • Context: supplier token, GL account, cost center or department, material group, plant or site, UoM, quantity and price band, currency, posting date.
  • Label: code, taxonomy name and version, level of granularity, and the full path (segment through commodity or brick).
  • Provenance: who coded it (buyer, supplier, rule, model, reviewer), whether it was reviewed, and the review date.

For upstream document capture (pulling the line items off a scanned invoice) see purchase orders licensed for AI training; this page covers the classification step that follows extraction.

UNSPSC, GS1 GPC and custom taxonomies compared

The target taxonomy determines code depth, label sources and how much crosswalk work a supplier's data will need. UNSPSC is a four-level hierarchy of segment, family, class and commodity encoded in eight digits, two per level, so a commodity code nests inside its class, family and segment codes. GS1 Global Product Classification uses segment, family, class and brick, with 44 segments in the June 2024 version described by GS1 Nederland [1]. Retail catalog teams also meet Google's google_product_category, which accepts either a numeric ID or a full path, while merchants put their own labels in product_type [2].

TaxonomyTypical homeDepth and formWhere labels usually come fromCommon sourcing problem
UNSPSCIndirect procurement, public sector, P2P suites4 levels, 8 digitsSupplier catalogs (often coarse), buyer coding, spend-analytics toolsCodes pinned to different versions; many lines coded only to segment or family
GS1 GPCRetail and CPG master data, GDSN item data4 levels to brick [1]Brand owners publishing item dataBuilt for traded products; thin on services and much indirect or MRO spend
Google product categoryE-commerce feedsFixed path or ID [2]Merchants and feed toolsMerchant product_type is free-form and inconsistent
Custom company taxonomyCategory management, sourcing strategyUsually 2 to 4 levelsCategory managers, rules enginesFrequent restructuring; needs a crosswalk to a standard

Most enterprise buyers end up needing both a standard code and the company's own category, because the operating decisions (which category manager owns a line, which sourcing event it belongs to) are made in the custom tree. Ask for the crosswalk table the supplying company used, not only the final codes.

Label provenance decides whether the labels are worth training on

Label provenance is the single biggest quality variable, because spend codes are assigned by very different processes with very different error profiles. A line coded by a category analyst during a spend cube refresh is a different object from a supplier's self-assigned catalog code or a code written by a regex rule in 2019. If provenance is mixed and unrecorded, your model learns the rules engine's mistakes as ground truth.

Annotation conflicts, where identical text carries different codes, are common wherever several departments or tools code independently, and they cap the accuracy any model can show against those labels. Ask suppliers for conflict rates per taxonomy level, and resolve conflicts or keep them as soft labels on purpose rather than dropping them silently.

Use a per-row label_source field and audit it before acceptance. The method in how to audit annotation quality in a labeled dataset applies directly: sample per label source, have an independent coder re-label, and compute agreement at each hierarchy level rather than only at the leaf.

Splits, duplicates and leakage in line-item data

Random splits overstate accuracy on spend data because recurring purchases produce the same text many times, so identical lines land in both train and test. This mirrors general findings that near-duplicates are widespread in training corpora and distort both training and evaluation [4].

Practical controls a buyer should require or apply on receipt:

  • Group by normalized text (lowercase, collapsed whitespace, stripped part numbers) and split by group, not by row.
  • Split by supplier or by company to test generalization to unseen vendors, which is what a new customer deployment looks like.
  • Split by time (train on earlier fiscal years, test on later) to expose taxonomy-version drift and new products.
  • Report metrics per level (segment, family, class, commodity) and spend-weighted as well as count-weighted, since a model can be accurate on counts and wrong on the dollars.

Taxonomy version drift and crosswalks

Codes are only comparable when the taxonomy version is recorded, because standard taxonomies are revised and custom trees are restructured during category reorganizations. UNSPSC releases add, retire and move codes; a model trained on a mix of versions learns contradictory mappings for the same item. Custom trees tend to change more often than standard ones.

Ask each supplying company for the version identifier attached to every coded row, the date of any re-coding project, and the mapping table used when the tree changed. Where the data spans multiple versions, either remap to a single target version with the mapping recorded or keep the version as a feature and evaluate per version. The same discipline applies to adjacent code systems such as tariff codes; see customs entry and HTS classification data.

Long-tail coverage requires many companies

Category coverage is the main reason to license from several organizations rather than one. A single company's spend concentrates in a few hundred commodities, while public product datasets show how wide the tail is: Google's MAVE spans 1,257 product categories [3]. MAVE and similar resources come from web product pages, not from procurement ledgers, so they do not reproduce terse AP text, GL context or services spend.

Plan coverage the way you plan a stratified sample. List the segments and families your deployment must handle, set a minimum labeled count per family, and ask suppliers for a code-frequency table before any agreement. Services lines (consulting, maintenance, facilities, freight) are under-represented in product-centric data and usually need deliberate sourcing.

Confidentiality: supplier names, prices and personal data

Supplier names and negotiated unit prices are commercially sensitive to the company that holds them, so expect them to be tokenized, bucketed or aggregated before release. A stable pseudonymous supplier token preserves most of the signal a classifier needs, while price bands (rather than exact prices) protect negotiated terms. P-card and expense lines can also carry cardholder names, employee IDs or traveler names in free text, which need removal before delivery; NIST's survey of de-identification shows that removed identifiers do not rule out re-identification from remaining fields [8].

Agree on these choices in the request, because tokenization changes what you can test. If supplier is a key feature, ask for consistent tokens across the whole delivery and across future refreshes.

Illustrative record and request template

A good delivery is a flat line-level table (Parquet or JSON Lines) with a data card describing sources, coding methods and intended use [5], ideally with machine-readable metadata such as Croissant [6].

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "line_id": "c07-po-000418-3",
  "company_token": "C07",
  "source_doc": "purchase_order",
  "erp": "SAP S/4HANA",
  "line_text": "GLOVE NITRILE PF L 100/BX",
  "long_text": null,
  "supplier_token": "S-5512",
  "gl_account_bucket": "lab_supplies",
  "material_group": "LABCONS",
  "uom": "BX",
  "price_band_usd": "10-25",
  "posting_month": "2025-03",
  "label": {
    "taxonomy": "UNSPSC",
    "taxonomy_version": "recorded by supplier",
    "level": "commodity",
    "code": "<8-digit code>",
    "custom_category": "Lab & Scientific > Consumables > PPE",
    "label_source": "category_analyst",
    "reviewed": true,
    "review_date": "2025-06-30"
  }
}

Request checklist for suppliers:

  1. Source systems and document types, with date range per company.
  2. Taxonomies used, version per row, and crosswalk tables.
  3. Code-frequency table by level, plus share of rows coded only to segment or family.
  4. label_source distribution and any re-coding history.
  5. Conflict rate for identical normalized text.
  6. Treatment of supplier names, prices and free-text personal details.
  7. Format, data card and refresh cadence.

How SourceX handles requests for labeled line items

SourceX sources operational datasets from US companies on request, including finance workflows where coded line items often live; categories describe what can be requested, not stock, and a request does not guarantee a match. You describe the data, and SourceX looks for US businesses that hold it; each release is approved by the supplying company. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery, with names, emails, phones and account numbers removed or replaced and the method recorded. You can describe your line-item and taxonomy requirements as a buyer.

Related reading: the tabular, time-series and transactional data guide, ERP transaction and master data for AI training, product categorization and taxonomy-mapping data for retail catalogs, structured product attribute data, the procurement records overview and procurement assistants use cases.

Request spend classification training data

SourceX sources operational datasets, including finance workflow records, from US companies for AI teams wherever they are based, and manages licensing from assessment through agreement and ongoing purchases. Nothing is contracted until a supplier agrees. Describe the line items, taxonomies and label provenance you need at sourcex.si/buyers.

Frequently asked questions

Can I train a spend classifier on public product data alone?

You can pretrain on it, but public product datasets come mostly from web catalogs [3] and lack AP and PO line text, GL accounts and services spend. Expect a gap on terse internal descriptions and on services categories, and hold out real procurement lines for evaluation.

Which taxonomy level should labels reach?

Ask for the deepest level the supplying company actually coded and record the level per row. Many organizations code reliably to family or class and only partially to commodity, so leaf-level accuracy claims on mixed-depth data are misleading.

Should I accept supplier-assigned catalog codes as labels?

Accept them as a separate label source with their own audit. Supplier-coded catalog codes reflect the seller's view and are often coarse, so compare them against buyer-reviewed codes on a sample before mixing them into training.

Sources

  1. GS1 Nederland, "GPC in a nutshell (June 2024)" (2024). https://www.gs1.nl/media/sfpiaxye/gpc-in-a-nutshell_jun24-def.pdf
  2. Google Merchant Center Help, "Google product category [google_product_category]". https://support.google.com/merchants/answer/6324436?hl=en-GB
  3. Google Research, "MAVE: A Product Dataset for Multi-source Attribute Value Extraction" (2022). https://research.google/pubs/mave-a-product-dataset-for-multi-source-attribute-value-extraction/
  4. Lee et al., ACL 2022, "Deduplicating Training Data Makes Language Models Better" (2022). https://arxiv.org/abs/2107.06499v1
  5. Pushkarna, Zaldivar, Kjartansson (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  6. Google Research, "Croissant: a metadata format for ML-ready datasets". https://research.google/blog/croissant-a-metadata-format-for-ml-ready-datasets/
  7. SAP Learning, "Releasing Purchase Requisitions". https://learning.sap.com/courses/purchasing-in-sap-s-4hana/releasing-purchase-requisitions
  8. NIST, "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data