Skip to content

Tables, time series and transactional data

Fraud-Labeled Transaction Data for Fraud Detection Models

Quick answer

Fraud detection training data is a stream of real card, ACH or e-commerce transactions joined to confirmed outcomes: chargebacks with reason codes, issuer-confirmed fraud, ACH unauthorized returns or analyst dispositions. It must also carry the timestamp at which each label became known. Public Kaggle sets such as IEEE-CIS are useful for prototyping, but they lack label timing, stable account histories and current attack patterns. Production-grade data comes from processors, issuers and merchants under license, with tokenized identifiers and a disclosed fraud base rate.

By SourceX Editorial · Updated

Why public fraud datasets stop working past prototyping

Public fraud sets run out at the point where a model has to match a live decision stream. One 2026 study concluded from its review of open sources that it had to build its own dataset, merging several public Kaggle fraud sets and adding synthetic attributes [1]. That kind of stitching produces columns that were never observed together. Interactions between device, merchant category and velocity features then do not reflect what any real fraud ring did.

Production fraud work starts from operating data instead. A vendor purchase-card fraud use case, for example, begins with a data feed set up with a bank or processing company, not a downloaded file [2]. Public sets also hide the two things that break models in deployment: when labels arrive and how the fraud rate moves over time.

Labels and label maturity: what a fraud outcome actually means

A fraud label is only usable if you know its source and when it was observed. Transactions are scored in milliseconds, but labels arrive in at least two waves. Analysts confirm a small number of alerts quickly, and most other labels arrive days or weeks later when cardholders dispute charges or banks return debits. Feedback from reviewed alerts is also biased toward whatever the incumbent model flagged, so keep it distinguishable from delayed dispute labels.

When you specify labels, ask for these fields per transaction:

  • label_source: chargeback, issuer fraud report (for example a TC40 or SAFE-style report where the holder has one), ACH return, manual review, or a rule hit (exclude rule hits from ground truth).
  • label_reason: network dispute reason code, or ACH return code. R10 (unauthorized) and R05 should be kept distinct from R01 insufficient funds, which is credit loss rather than fraud [5].
  • label_observed_at: the date the outcome was recorded, so you can rebuild what the model would have known at decision time.
  • maturity_window: the cutoff after which unlabeled rows are treated as legitimate. Rows newer than that window should be flagged as immature, not counted as negatives.

Without label_observed_at, temporal splits leak future knowledge and offline recall looks better than it will in production. Friendly fraud (first-party disputes) also needs its own code. Otherwise a model trained on chargebacks learns to flag customers who dispute, not stolen credentials. For dispute workflows themselves, see chargeback representment case data.

Base rates, sampling and the evaluation set

Disclose the natural fraud rate and keep at least one evaluation slice at that rate. Suppliers often down-sample legitimate traffic to keep files small. That is acceptable for training if the sampling rate per stratum is documented, but a balanced test set inflates precision and makes PR-AUC numbers incomparable across vendors. Report the monthly base rate too, because a fraud rate that moves with attack waves changes threshold settings.

Label quality matters as much as volume. Test-set label errors can change which model ranks first on a benchmark [6], and fraud labels are noisy by construction: unreported fraud sits in the negatives and friendly fraud sits in the positives. Ask how many confirmed cases were re-reviewed and whether disputes reversed in representment were relabeled. The outcome-labeled evaluation data guide covers holdout design in more depth.

Account histories, perspective and drift

Pattern features need per-account histories, not isolated rows. Card-fraud approaches built on transaction patterns depend on each card's prior activity [3]. In practice that means stable tokens for card, account, device and merchant that persist across every delivery and refresh. Velocity, geo-jump and new-merchant features only work if the same customer keeps the same token over time.

Perspective decides which model the data can train:

PerspectiveTypical fieldsModels it trainsCommon gap
IssuerAuthorization messages, MCC, POS entry mode, decline codes, cardholder disputesAuthorization scoring, account takeoverNo basket or device detail
Acquirer or processorCross-merchant authorizations, settlement, chargebacksMerchant risk, BIN attack detectionLabels depend on issuer reporting
Merchant or e-commerceOrders, basket, device fingerprint, shipping vs billing, login eventsCheckout fraud, promo abuse, account takeoverMostly chargeback labels, seen late
ACH originatorEntry class (WEB, PPD), return codes, account ageUnauthorized debit, first-party fraudReturn windows differ by code [5]

Drift is the other reason to license recent data. Customer habits and fraud strategies both change over time, so ask for a continuous span of at least one full seasonal cycle and for ongoing refreshes on the same schema. The general transaction licensing questions are covered on our financial transaction data page. Ledger-side anomalies belong with journal entry anomaly detection data, and altered invoices with document fraud detection data.

Request template for a fraud-labeled transaction license

Illustrative example: invented to show structure; it does not describe an available dataset.

request: fraud-labeled card-not-present transactions
perspective: merchant (US e-commerce, digital goods and physical goods)
grain: one row per authorization attempt, including declines
span: 24 consecutive months, plus monthly refresh on the same schema
identifiers:
  card_token: keyed hash or index token, stable across all deliveries
  account_token, device_token, merchant_id: stable tokens
  pan: never delivered
fields: [auth_ts, amount, currency, mcc, entry_mode, avs_result, cvv_result,
         bin_country, ip_country, ship_bill_match, decision, decline_code]
labels:
  label_source: [chargeback, issuer_fraud_report, manual_review]
  label_reason: network reason code; friendly fraud coded separately
  label_observed_at: required
  maturity_window_days: 120
sampling: natural rate for eval slice; documented per-stratum rates otherwise
documentation: data dictionary, base rate by month, known pipeline changes

Privacy, PCI scope and reuse rights

Primary account numbers should never be part of a training delivery. PCI DSS Requirement 3 requires stored PAN to be rendered unreadable, through methods such as keyed hashing of the full PAN, truncation, index tokens or strong cryptography [4]. An unkeyed hash is weak because PANs are enumerable. Ask the supplier to tokenize with a secret it keeps, so tokens stay consistent across deliveries but cannot be reversed by the buyer. Our guide on de-identifying financial transaction data covers names, addresses and free-text fields.

Rights matter as much as masking. Under the FTC's GLBA guidance, a business that receives nonpublic personal information outside an exception steps into the shoes of the originating institution, which limits how that data can be reused and redisclosed [8]. Confirm the supplier has the right to share the data for model training before it moves. SourceX rights-reviews each dataset for ownership and consents and delivers it under a license that defines records, uses, term and delivery.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Real vs synthetic fraud data

Synthetic fraud data is best treated as augmentation, measured against a real reference set. Tools such as SDMetrics score a synthetic table against the real data it was modeled on [7]. That means you still need a real sample to validate it. Synthetic generators also reproduce the attack patterns they were given. They cannot supply the next fraud pattern, which is the main reason to license fresh labeled data. For tabular pretraining mixes, see tabular foundation model data.

How SourceX handles fraud-labeled transaction requests

SourceX sources operational datasets from US companies and manages the commercial process, including licensing and ongoing purchases. Data is sourced on request rather than held in stock, so a request does not guarantee a match. Every release must be approved by the supplying company. Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement. Start by describing the data, not the businesses, on the SourceX buyer page. See the structured data buyer's guide for related table types.

License fraud detection training data with confirmed labels

SourceX looks for US businesses that hold the transaction data you describe and handles assessment, pricing and allowed uses in a license. Nothing is contracted until a supplier agrees, and terms are set per deal. Describe the fraud-labeled data you need.

Frequently asked questions

Is the IEEE-CIS dataset enough to benchmark a production fraud model?

It is useful for comparing methods, but it has no label-observation dates and its anonymized columns cannot be matched to your own features. Researchers who need more realistic conditions have built merged datasets with synthetic attributes instead [1].

How long should the label-maturity window be?

Set it from the dispute and return windows for your rails, then confirm it from the supplier's own data by plotting cumulative labels against transaction age. Treat rows younger than the window as unlabeled.

Should declined transactions be included?

Yes, if the goal is authorization scoring. Excluding declines hides the attacks the incumbent system already stopped and biases the model toward the fraud that got through.

Sources

  1. Information Technology and Mathematical Modelling journal, "Methodology of dataset preparation for training e-commerce fraud detection models" (2026). https://journals.nmetau.edu.ua/index.php/itmm/en/article/view/2468
  2. DataRobot, "Purchase card fraud detection (business accelerator)". https://docs.datarobot.com/en/more-info/biz-accelerators/p-card-detect.html
  3. arXiv, "A data mining approach using transaction patterns for card fraud detection" (2013). https://arxiv.org/pdf/1306.5547
  4. HeroDevs, "PCI DSS 4.0 Requirement 3: How to Protect Stored Account Data". https://www.herodevs.com/blog-posts/pci-dss-4-0-requirement-3-how-to-protect-stored-account-data
  5. Plaid, "ACH return codes: A complete guide for businesses". https://plaid.com/en-eu/resources/ach/ach-return/
  6. Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  7. DataCebo / Synthetic Data Vault, "SDMetrics". https://docs.sdv.dev/sdmetrics
  8. Federal Trade Commission, "How To Comply with the Privacy of Consumer Financial Information Rule of the Gramm-Leach-Bliley Act". https://www.ftc.gov/business-guidance/resources/how-comply-privacy-consumer-financial-information-rule-gramm-leach-bliley-act

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data