Skip to content

Industry-specific operational data

Transaction enrichment data: merchant name normalization and categorization labels

Quick answer

Transaction enrichment training data pairs raw bank or card descriptors, such as SQ BLUE FERN CAFE 0412 PORTLAND OR, with a verified merchant entity, a category from your taxonomy and, ideally, recurrence and counterparty labels. MCCs and user recategorizations are useful but noisy label sources; a human-verified gold set stratified by merchant frequency is what makes evaluation trustworthy. Before licensing app-collected descriptors or corrections, confirm the original terms covered model training and whether bank data arrived through an aggregator.

By SourceX Editorial · Updated

What a raw descriptor contains and what normalization must recover

A raw descriptor is a lossy, processor-shaped string, and normalization must recover a stable merchant identity from it. The same coffee shop can appear as SQ *BLUE FERN CAFE, SQ *BLUEFERN CAF 0412, POS DEBIT 10/03 BLUE FERN CAFE PORTLAND or, after an ACH debit, as an originator name the consumer never saw at the counter. Transaction-pattern models, such as card fraud detectors, build features from a customer's transaction history [8], and those features, like recurring-bill detection and merchant-level spend, assume one entity per merchant.

The noise has recognizable layers, and a useful dataset keeps each one visible rather than pre-cleaning it:

  • Processor and wallet prefixes: payment-facilitator and wallet tokens (SQ *, TST*, PAYPAL *, APPLE.COM/BILL) that name the intermediary, not the merchant.
  • Truncation: card networks and core banking systems cap descriptor length, so INTERNATIONAL HOUSE OF and INTL HOUSE OF PANC must resolve to the same chain.
  • Store and terminal identifiers: store numbers, terminal IDs and location codes (#0412, STORE 2231) that matter for location models but must not split the entity.
  • Location fragments: city, state or phone fragments appended by the acquirer, sometimes the merchant's registered office rather than the point of sale.
  • Rail-specific formats: ACH entries carry an originator name and entry description, wires carry free-text remittance, and P2P transfers carry a counterparty handle, each with different failure modes.

For buyers, the practical question is whether the supplier kept the raw string exactly as received from the core, processor or aggregator feed. A dataset whose "raw" column was already passed through a vendor enrichment API teaches your model to imitate that vendor, including its errors.

MCCs versus app categories: two different label spaces

A merchant category code is an acquirer-assigned business classification, not a spending category, so treat it as a feature or a weak prior rather than ground truth. ISO 18245 defines MCC values that classify merchants by type of business, trade or service [1]. The acquirer assigns the code at onboarding, which means a supercenter selling groceries, fuel and pharmacy items carries one code for every basket, and a miscoded merchant stays miscoded until someone notices.

Consumer-facing categories differ by app. A PFM app may split "Groceries" from "Household", an accounting platform maps to a chart of accounts, and a lender's affordability model may care only about essential versus discretionary spend. An MCC-to-category mapping table is therefore a product decision, and a training set should carry the supplier's own taxonomy with its version, plus the mapping rules in force when labels were assigned.

Ask suppliers which field the label came from: the MCC in the authorization message, a vendor enrichment category, a rules engine, a user correction or a human reviewer. Each has a different error profile, and mixing them silently is a frequent reason a categorization model plateaus.

User recategorizations as weak labels

User recategorizations are abundant and cheap but biased, so use them for training signal and keep them out of evaluation. A user who moves a transaction from "Shopping" to "Gifts" is expressing a personal budget preference, not correcting the merchant's identity. Corrections also cluster on the transactions users look at, typically large or unusual ones, so the long tail of small descriptors is under-corrected.

Treat a correction event as a record with its own fields: the prior label and its source, the new label, the timestamp, whether the user applied a "remember this merchant" rule, and how many distinct users made the same change. Agreement across users on the same normalized merchant is a far stronger signal than any single edit. Label noise matters even in curated sets: an audit of widely used benchmarks estimated an average test-set label error rate of at least 3.3% [7], and our guide to finding label errors with confident learning covers how to triage noisy fields before training.

Building a merchant entity resolution gold set

Merchant entity resolution needs its own gold set because public entity-resolution benchmarks are built from other domains. Well-known benchmarks with gold match labels, such as WDC Block, come from product records [5]; none model processor prefixes, truncation or ACH originator names. Evaluation should also be entity-centric, scoring how well clusters match true entities rather than only pairwise match accuracy [6].

A practical gold set has three properties. It is stratified by merchant frequency, because head merchants are easy and new or long-tail merchants drive most errors. It is double-annotated with an adjudication step and recorded disagreement. And it is frozen with a time window, so you can measure drift as new merchants, rebrands and processor changes appear. For a broader view of using operational outcomes as ground truth, see outcome-labeled evaluation data and verifying ground truth in operational records.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueNotes for buyers
txn_idt_8f31c2Pseudonymous; stable within the delivery
account_tokenacct_19abReplaces account number; supports per-customer history
railcard_debitcard_debit, card_credit, ach, wire, p2p
raw_descriptorSQ *BLUEFERN CAF 0412 PORTLAND ORExactly as received; no pre-cleaning
amount, posted_date6.75, 2026-03-04Date shifting may be applied; ask how
mcc5814From authorization data, when available
merchant_entity_idm_bluefernResolved entity; cluster ID across descriptors
merchant_display_nameBlue Fern CafeHuman-verified for gold rows
category_label, taxonomy_versiondining.coffee, v7Supplier taxonomy plus version
label_sourcehuman_goldmcc_map, vendor, rules, user_correction, human_gold
recurring_flag, recurrence_periodfalse, nullFor subscription and bill detection
user_correction_count0Distinct users who changed the label
merchant_frequency_bucketlong_tailhead, torso, long_tail, new

Self-generated data versus consumer-permissioned bank data

Who generated the data determines what you can license, so separate first-party records from consumer-permissioned data pulled through aggregators. A fintech's own support tickets, app event logs, category rules and user corrections are records the company created about its service. Descriptors pulled from a consumer's bank account through an aggregator arrive under the aggregator's contract, the consumer's authorization and, for financial institutions, GLBA's Regulation P limits on reuse and redisclosure of nonpublic personal information [4].

The CFPB's personal financial data rights rule under Section 1033 of the Dodd-Frank Act addresses how authorized third parties may use permissioned data, and the final rule was issued in October 2024 with compliance for large institutions beginning in 2026 [9]. Check its status as of October 2026 with counsel before relying on any reading. The deeper legal analysis of de-identified financial data belongs on our de-identification resource, how to de-identify financial transaction data.

Consent at collection is the other gate. FTC staff warned that quietly or retroactively changing terms to use consumer data for AI training can be unfair or deceptive [2], and that companies can be liable for breaking promises not to use customer data for undisclosed purposes such as model training [3]. For descriptors and corrections collected inside an app, ask for the terms version in force at collection and whether it covered training use by the company and licensing to third parties.

Buyer checklist for transaction enrichment datasets

A short, specific request filters out mismatched supply faster than a long questionnaire. Use this checklist when scoping a request or reviewing a sample.

Illustrative example: invented to show structure; it does not describe an available dataset.

  1. Rails and feeds: which rails (card, ACH, wire, P2P), which feeds (core, processor, aggregator), and whether raw strings are untouched.
  2. Label provenance: per-row label_source, taxonomy version, MCC availability and the MCC-to-category mapping in force.
  3. Entity resolution: a merchant entity ID that clusters descriptors, with rebrand and acquisition handling documented.
  4. Gold subset: size, frequency stratification, double annotation, adjudication rules and inter-annotator agreement.
  5. Corrections: prior and new labels, timestamps, rule creation and distinct-user counts.
  6. Recurrence: recurring flags, period and amount tolerance used to assign them.
  7. Privacy treatment: how account numbers, names in P2P memos and payroll descriptors were removed or replaced, and whether merchant names that are actually individuals (sole proprietors, landlords) were handled.
  8. Consent and contracts: terms version at collection, aggregator contract restrictions and any Regulation P constraints.
  9. Leakage controls: time-based splits so new merchants in the test window are truly unseen.

Related operational datasets follow the same pattern: product categorization and taxonomy-mapping data for catalog taxonomies, merchant onboarding and KYB underwriting decisions for the acquirer side, and bank statements and income documents for document-based cash-flow models. The industry-specific operational data hub lists the rest.

How SourceX approaches descriptor and categorization requests

SourceX sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases. Data is sourced on request rather than held in stock, so a request for descriptor and label data starts a search for US businesses that hold it, and a request does not guarantee a match. Buyers describe the data, not the businesses, and every release is approved by the supplying company.

Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. SourceX does not source scraped web content or standalone contact lists. Fintech buyers can review the fintech software buyer page, the broader financial transaction data licensing page, accounting reconciliation datasets for GL coding, or start at SourceX for AI data buyers.

Source transaction enrichment training data

SourceX finds US companies that hold operational data, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is transacted. Nothing is contracted until a supplier agrees. Describe the descriptors, labels and rails you need at SourceX for AI data buyers.

Frequently asked questions

Can I train a categorization model on MCCs alone?

You can train a baseline, but MCCs describe the merchant's business type as assigned by the acquirer, not the purpose of a given purchase [1]. Multi-category retailers, miscoded merchants and non-card rails without MCCs cap accuracy, so pair MCCs with descriptor text and human-verified labels.

How large should a merchant normalization gold set be?

There is no universal number; size it so each frequency bucket (head, torso, long tail, new) has enough resolved entities to report per-bucket precision and recall with useful confidence intervals. Long-tail and new merchants usually need deliberate oversampling.

Are user corrections safe to use as evaluation labels?

Generally no. They reflect personal budgeting choices and are concentrated on salient transactions, so use them as weak training signal and evaluate against an adjudicated human gold set. This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Sources

  1. International Organization for Standardization, "ISO 18245:2023 Retail financial services - Merchant category codes" (2023). https://www.iso.org/standard/79450.html?browse=tc
  2. Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  3. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  4. Consumer Financial Protection Bureau, "Regulation P: Privacy of Consumer Financial Information (12 CFR Part 1016)". https://www.consumerfinance.gov/rules-policy/regulations/1016/1/
  5. University of Mannheim, Data and Web Science Group, "WDC Block: A large Blocking Benchmark released". https://www.uni-mannheim.de/dws/news/wdc-block-a-large-blocking-benchmark-released/
  6. arXiv, "How to Evaluate Entity Resolution Systems: An Entity-Centric Framework with Application to Inventor Name Disambiguation" (2024). https://arxiv.org/pdf/2404.05622
  7. Northcutt, Athalye, Mueller (arXiv / NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  8. arXiv, "A data mining approach using transaction patterns for card fraud detection" (2013). https://arxiv.org/pdf/1306.5547
  9. Consumer Financial Protection Bureau, "Personal Financial Data Rights (Section 1033)" (2024). https://www.consumerfinance.gov/rules-policy/final-rules/personal-financial-data-rights/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data