Skip to content

Industry-specific operational data

Merchant onboarding and KYB underwriting decisions for payments risk models

Quick answer

Merchant underwriting data for AI is the record of how a payment facilitator, acquirer or embedded-payments platform decided whether to board a business: the application, KYB evidence on the entity and its beneficial owners, website and product review notes, the assigned merchant category code (MCC), the approve, decline or conditional decision with reserves and limits, and what happened after boarding. The outcome link (chargebacks, monitoring-program enrollment, termination) is what turns an onboarding archive into training and evaluation data.

By SourceX Editorial · Updated

What a merchant underwriting record must contain to be useful

A usable record joins four layers that usually live in different systems: the application, the evidence gathered, the decision and the post-boarding outcome. Applications typically sit in an onboarding or CRM tool, KYB vendor responses in API logs, reviewer notes in a case-management queue, and outcomes in the risk or settlement platform. A file that has only the application and the decision trains a model to imitate past reviewers; it cannot tell you whether those reviewers were right.

The core fields buyers should ask for:

  • Application: legal name (pseudonymized), entity type, state of formation, years in business, stated products, processing volume and average ticket estimates, card-present versus card-not-present mix, refund and delivery-time policy, prior processor history.
  • KYB evidence: registry match results, EIN or tax-ID verification status, beneficial-owner identification outcome, sanctions and watchlist screening result codes, address and phone verification, document types collected (articles, bank letter, processing statements).
  • Review artifacts: website review notes (product pages, terms, refund policy, checkout flow), prohibited and restricted business screening, MCC assignment and any later MCC change, reviewer identity as a stable pseudonymous ID, escalation path.
  • Decision: approve, decline, approve with conditions, or pend for information; reserve type and percentage, rolling or fixed holds, volume caps, delayed funding; reason codes and free-text rationale.
  • Outcomes: chargeback and fraud ratios by month after boarding, card-network monitoring-program enrollment, account closure reason, and any placement on a network terminated-merchant list such as Mastercard's MATCH.

For beneficial ownership, US covered institutions follow FinCEN's customer due diligence rule at 31 CFR 1010.230, which as of October 2026 includes risk-based relief for repeat verification under the 2026 exceptive relief order [7]. Payment facilitators sit under a sponsor bank's program, so their KYB fields often mirror that bank's CDD requirements. Expect owner-verification fields to change shape around the rule's May 2018 compliance date and later relief in any multi-year archive.

MCC classification data and why its labels are noisy

MCC assignments are useful classification labels but should be treated as noisy, not as ground truth. MCCs are four-digit codes whose values are defined in ISO 18245:2023 [5]. Card networks publish their own implementations and rules on top of that standard, and an acquirer's assigned MCC reflects both the network list and the reviewer's reading of the business.

Common error patterns in MCC labels:

  • Default codes: onboarding forms that pre-select a generic retail or services code when the applicant skips the field.
  • Blended businesses: a merchant selling both supplements and fitness apparel gets whichever code the reviewer picked first.
  • Drift: the business pivots after boarding, but the MCC is never updated.
  • Downgrading: a reviewer, or the applicant, chooses a lower-risk or lower-interchange code than the products justify.

Label errors matter even at small rates: an audit of widely used benchmarks estimated an average test-set label error rate of at least 3.3%, enough to reorder model rankings [1]. Confident learning, implemented in the open-source cleanlab package, ranks likely mislabeled examples by estimating class-conditional noise [2]. Ask whether the supplier can provide MCC change history so you can separate initial assignment from corrected assignment; the corrected value is usually the better label. For adjacent normalization work on transaction streams, see transaction enrichment and merchant normalization data.

Outcome labels for merchant risk scoring

The strongest merchant risk labels are post-boarding outcomes measured over a fixed window, not the onboarding decision itself. Declined applicants have no outcome, so a model trained on approvals only learns the risk of merchants the old policy already accepted. This is the classic reject-inference problem, and it should shape your evaluation design before you sign anything.

Practical outcome definitions buyers use:

LabelDefinition to requestFailure mode
Early lossNegative balance or reserve shortfall within 180 days of first transactionSmall sample; depends on reserve policy
Dispute riskMonthly chargeback count and amount ratios, months 1 to 12Mixes fraud and service disputes; see reason codes
Monitoring-program enrollmentCard-network excessive-dispute or fraud program flag with entry dateProgram rules change; timestamp each flag
TerminationAccount closed by the provider, with closure reason codeVoluntary churn mislabeled as risk closure
Terminated-merchant list placementNetwork list placement and reason codeRare; strongly provider-policy dependent

Card-network monitoring programs are revised over time, including Visa's consolidation of its dispute and fraud monitoring programs into the Visa Acquirer Monitoring Program (VAMP) in 2025 [6], so a flag from 2023 and a flag from 2026 may not mean the same threshold. Record the program name and version with each label. Chargeback outcomes pair naturally with chargeback representment case data if you also want to model dispute handling. For a method to test whether outcome fields are reliable, read verifying outcome labels in operational records.

Website review notes and review-agent training data

Review notes are the richest input for an onboarding review agent because they capture what a human checked and why. A good note says the reviewer found no refund policy, the checkout used a third-party domain, and the product pages listed a restricted ingredient; a weak note says "looks fine." Ask for a sample of 50 notes before licensing and score them for specificity.

For agent training, the ideal unit is a case: the application, the evidence the agent would have had at decision time, the reviewer's notes in order, the decision and the outcome. Snapshots of the merchant's website at review time are valuable but often were not retained, or were retained only as screenshots. Confirm whether the supplier holds the captured pages or only URLs, because a URL fetched today shows a different site. If you also need related analyst work, fraud investigation case notes and analyst decisions and KYC and CDD case review files cover individual-customer and investigation workflows rather than merchant boarding.

Privacy, rights and de-identification for KYB data

KYB files mix business data with personal data about owners and signers, and the personal layer has to come out before licensing. Owner names, dates of birth, Social Security numbers, ID document images, home addresses and personal phone numbers should be removed or replaced; business identifiers such as legal name, EIN and domain can be pseudonymized with a stable token so cases still join across tables. Sole proprietors are a trap: the business name is often the owner's name, and the tax ID on file may be the owner's SSN rather than an EIN.

Gramm-Leach-Bliley privacy rules protect consumers, so business-purpose merchant applications often fall outside them, but any consumer data in the file needs checking. The FTC notes that a recipient of nonpublic personal information outside a Gramm-Leach-Bliley exception steps into the shoes of the originating institution for reuse and redisclosure [3]. Ask the supplier's counsel whether any consumer information is included, which exception a transfer would rely on, and whether the sponsor bank's program agreement permits it. Bias review also matters, since past decline patterns can encode geography or ownership proxies; see historical decision bias in operational labels.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

A request template for merchant underwriting data

Write the request as a dataset specification, not a vendor wish list, so suppliers can answer yes or no on each line. Document the result in a datasheet or data card covering sources, collection, labeling and known limits [4].

Illustrative example: invented to show structure; it does not describe an available dataset.

request: merchant_onboarding_decisions
unit: one onboarding case (application through 12-month outcome)
population:
  provider_type: payment facilitator or acquirer, US merchants
  period: 2022-01 to 2025-12 applications
  include_declines: true          # needed for reject inference
fields_required:
  - case_id, applied_at, decided_at
  - entity_type, state, years_in_business (bucketed)
  - stated_products_text, website_domain_token
  - kyb_checks: [registry_match, tin_match, sanctions_result, ubo_verified]
  - mcc_initial, mcc_final, mcc_changed_at
  - decision: [approve, decline, conditional, pend]
  - reserve_type, reserve_pct, volume_cap
  - reviewer_id_token, review_notes_text
outcomes_required:
  - monthly_dispute_ratio m1..m12
  - monitoring_program_flag + program_name + flag_date
  - closure_reason_code, closed_at
removed_before_delivery:
  - owner names, DOB, SSN, ID images, home address, personal phone
documentation: data card, field dictionary, de-identification method

How SourceX handles merchant underwriting requests

SourceX sources operational datasets from US companies on request; it does not hold this data in stock, and a request does not guarantee a match. Buyers describe the data they need, SourceX looks for US businesses that hold it, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery.

Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Payments teams can review the fintech software buyer page, the overview of what AI companies build with payment processors data and the page for licensing underwriting files, or start a request on the SourceX buyers page.

Request merchant underwriting data for AI

SourceX sources operational datasets from US companies on request, which can include merchant applications, review records and decision histories when a supplier holds them, and manages licensing and ongoing purchases. Nothing is contracted until a supplier agrees, and terms are set per deal. Describe the cases, fields and outcomes you need on the SourceX buyers page.

Sources

  1. Northcutt, Athalye, Mueller (NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  2. Northcutt, Jiang, Chuang (arXiv), "Confident Learning: Estimating Uncertainty in Dataset Labels" (2019). https://arxiv.org/pdf/1911.00068
  3. Federal Trade Commission, "How To Comply with the Privacy of Consumer Financial Information Rule of the Gramm-Leach-Bliley Act". https://www.ftc.gov/business-guidance/resources/how-comply-privacy-consumer-financial-information-rule-gramm-leach-bliley-act
  4. Pushkarna, Zaldivar, Kjartansson (Google Research, FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  5. ISO 18245:2023, "Retail financial services — Merchant category codes" (2023). https://www.iso.org/standard/79450.html
  6. Visa, "Visa Acquirer Monitoring Program (VAMP) Fact Sheet 2025" (2025). https://corporate.visa.com/content/dam/VCOM/corporate/visa-perspectives/security-and-trust/documents/visa-acquirer-monitoring-program-fact-sheet-2025.pdf
  7. FinCEN, "FIN-2026-R001: Account Opening Exceptive Relief Order" (2026). https://www.fincen.gov/resources/statutes-and-regulations/cdd-rule-faqs

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data