Skip to content

Industry-specific operational data

Chargeback representment case data: reason codes, evidence and win/loss outcomes

Quick answer

Chargeback data for AI is most useful as complete merchant- or acquirer-side case files: the network reason code, the transaction and order record, the evidence packet the merchant submitted, the representment narrative, and the final outcome at each stage (first chargeback, pre-arbitration, arbitration) with recovered amount. Those joined records train representment drafting, evidence selection, win-probability and friendly-fraud models. Open data rarely contains them, so they are usually licensed from merchants, processors or chargeback-management operators, tokenized for PCI and de-identified.

By SourceX Editorial · Updated

What a representment case file contains

A usable representment case is a timeline of one disputed transaction joined to the evidence and decision at every stage, not a single fraud label. Fraud-labeled transaction tables, covered in fraud-labeled transaction data for fraud detection models, tell you whether a payment was fraudulent. A representment file tells you what the merchant argued, what proof it attached and whether the network outcome went its way.

Expect these record groups in a production case management system or acquirer dispute portal export:

  • Dispute header: network, reason code and description, dispute amount and currency, case and stage identifiers, received and due dates.
  • Transaction and order data: tokenized card reference, authorization result, AVS and CVV match codes, 3-D Secure status, order line items, merchant descriptor, MCC.
  • Fulfillment proof: carrier, tracking events, delivery confirmation, digital access or login logs for digital goods, service usage records.
  • Customer communications: support tickets, cancellation and refund requests, terms-of-service acceptance records.
  • Representment package: the narrative letter, the list of attached exhibits and the exhibit files (PDF, image, CSV of logs).
  • Outcome chain: representment accepted or rejected, pre-arbitration, arbitration filing, final liability and amount recovered or lost.

How network reason codes shape the labels

Reason codes are the primary stratification key, because evidence that wins one code is irrelevant to another. Visa groups dispute conditions into fraud (10.x), authorization (11.x), processing errors (12.x) and consumer disputes (13.x); for example 10.4 is card-absent fraud and 13.1 is merchandise or services not received. Mastercard numbers its reason codes differently, so a multi-network dataset needs a mapping table to a shared taxonomy. Network rules change, so check the current Visa and Mastercard rule books rather than vendor tables, which do not always agree with each other.

Treat code definitions as versioned labels. A case coded under a retired or merged code should keep its original code plus a normalized one, the problem described in code-set revisions in multi-year operational datasets.

Compelling Evidence 3.0 fields are high-value labels

Visa Compelling Evidence 3.0 gives friendly-fraud models a concrete, rule-defined target. Under CE 3.0, effective April 2023, a merchant can rebut a 10.4 card-absent fraud dispute by showing prior undisputed transactions from the same card at the same merchant. Visa's merchant guidance describes two prior transactions between 120 and 365 days old, with at least two matching data elements among user ID, IP address, shipping address and device ID, one of which must be the IP address or device ID [7].

For a buyer, that means the valuable fields are the ones merchants often fail to retain: device fingerprints, IP addresses at checkout, account IDs and the linkage to historical orders. Ask whether the source retained them per order and whether qualifying CE 3.0 submissions are flagged.

Merchant mix and stage coverage drive model validity

Reason-code distributions and win rates depend heavily on vertical, so a dataset from one merchant type will mislead a model deployed on another. Digital goods skew toward card-absent fraud and access-log evidence, travel toward cancellation and not-as-described disputes, and subscriptions toward canceled recurring and credit-not-processed claims. Record MCC and business model on every case so you can stratify splits and weight training.

Stage coverage matters as much. Many exports stop at the first representment decision, which leaves pre-arbitration and arbitration outcomes missing and biases win-probability labels toward early results. Apply the checks in case record completeness checks to confirm each case reaches a terminal outcome or is marked open.

Separate merchant-side files from issuer investigations

Representment files are the merchant and acquirer view; issuer dispute investigations are a different record type with different law. Card issuers resolve cardholder billing errors under Regulation Z [3], and financial institutions handle electronic fund transfer errors under Regulation E [2]. If you need the issuer side, see Reg E and Reg Z dispute investigation records.

Analyst reasoning from fraud operations also differs from representment narratives. For investigator notes and decisions, see fraud investigation case notes and analyst decisions.

PCI and privacy handling before licensing

Chargeback evidence is full of cardholder data, so PCI DSS scope and de-identification have to be resolved before any file leaves the source. PCI DSS Requirement 3 says to store account data only when needed, render stored PAN unreadable through truncation, tokens, keyed hashes or strong cryptography, and never store sensitive authentication data such as CVV after authorization; check the current PCI DSS text from the PCI Security Standards Council. A training dataset should contain no PAN at all: replace it with a consistent token so repeat-card features like CE 3.0 matching still work.

The harder part is unstructured evidence. Screenshots of order pages, shipping labels, emails and narrative letters carry names, street addresses, phone numbers and partial card numbers. Ask how redaction was applied to PDFs and images, not just to tables, and whether a reviewed sample exists; redaction training data covers the failure modes. IP addresses and device IDs are personal data in many regimes, so agree a hashing or salting method that preserves equality matching.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Case record schema and buyer checklist

A one-record-per-case JSON Lines file, with exhibits referenced by ID, keeps cases intact for SFT and outcome modeling [6]. Document it with a data card covering sources, collection, preparation and intended use [5].

Illustrative example: invented to show structure; it does not describe an available dataset.

{"case_id":"c_000184","network":"visa","reason_code":"10.4","reason_code_normalized":"fraud_card_absent","mcc":"5815","amount":49.99,"currency":"USD","card_token":"tok_9f2a","auth":{"avs":"Y","cvv":"M","three_ds":"frictionless"},"ce3":{"prior_txn_count":2,"matched_elements":["ip_address","user_id"],"submitted":true},"evidence":[{"exhibit_id":"e1","type":"access_log","format":"csv"},{"exhibit_id":"e2","type":"prior_orders","format":"pdf","redacted":true}],"narrative":"[REDACTED_NAME] logged in from the same device used for two prior undisputed orders...","stages":[{"stage":"first_chargeback","date":"2025-03-02"},{"stage":"representment","date":"2025-03-20","result":"won"}],"final_outcome":"merchant_won","amount_recovered":49.99}

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckWhy it mattersRed flag
Original and normalized reason codesCodes change across rule releasesOnly free-text reason labels
Terminal outcome per caseWin-probability labels need final liabilityCases end at "submitted"
Exhibit files, not just exhibit namesEvidence-selection models need contentNarratives without attachments
CE 3.0 inputs (IP, device, account ID)Friendly-fraud featuresDevice data dropped at export
MCC and business modelVertical shift changes distributionsSingle merchant, unstated vertical
PAN tokenization and image redactionPCI scope and privacyMasked PAN visible in screenshots
Label noise reviewDisputes are miscoded by cardholders and issuersNo disagreement or rework flags

For noisy labels such as a "fraud" code on what was a delivery complaint, confident learning can surface likely mislabeled cases for review [4].

Where licensed chargeback data comes from

Production dispute data reaches model builders through data agreements with merchants, processors and banks [1]. Expect to describe the record type, networks, verticals, stages and time window you need, then negotiate access with holders that can approve the release. Related buyer context is in the industry-specific operational data guide, fintech software buyers, finance buyers, licensing financial transaction data and what AI companies build with payment processor data. You can also describe a chargeback dataset request to SourceX.

Request chargeback representment data for AI

SourceX sources operational datasets from US companies on request, looking for businesses that hold the case records you describe; a request does not guarantee a match, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, with personal details such as names and account numbers removed or replaced before delivery. Describe the chargeback case data you need.

Sources

  1. DataRobot documentation, "Purchase card fraud detection". https://docs.datarobot.com/en/more-info/biz-accelerators/p-card-detect.html
  2. Electronic Code of Federal Regulations, "12 CFR 1005.11 Procedures for resolving errors". https://www.ecfr.gov/current/title-12/chapter-X/part-1005/subpart-A/section-1005.11
  3. Consumer Financial Protection Bureau, "12 CFR 1026.13 Billing error resolution". https://www.consumerfinance.gov/rules-policy/regulations/1026/13/
  4. Northcutt, Jiang and Chuang, "Confident Learning: Estimating Uncertainty in Dataset Labels" (2019). https://arxiv.org/pdf/1911.00068
  5. Pushkarna, Zaldivar and Kjartansson, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  6. jsonlines.org, "JSON Lines". https://jsonlines.org/
  7. Visa, "Evolution of Compelling Evidence – Merchant FAQs" (2023). https://usa.visa.com/content/dam/VCOM/regional/na/us/support-legal/documents/evolution-of-compelling-evidence-merchant-faqs-mar2023.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data