Industry-specific operational data
Fraud investigation case notes and analyst decisions for fraud-operations copilots
Quick answer
A fraud case notes dataset for copilot work is the investigation layer that sits on top of transactions: the alert that opened the case, each analyst action, customer-contact notes, the assigned typology, the decision and the loss and recovery outcome. Public fraud datasets contain labeled transactions but no case narratives, so this data has to be licensed from institutions that run fraud operations. Buyers should require account-level linkage, consistent typology codes, exclusion of anything that could reveal a SAR, and documented redaction of victim narratives.
By SourceX Editorial · Updated
Why fraud-labeled transactions are not enough for investigator copilots
Transaction tables teach a model to score risk, but an investigator copilot has to summarize a case, propose the next action and draft a rationale, which only case work can teach. Researchers already struggle with transaction data alone: a 2026 study concluded that existing open sources were insufficient and built its own set by merging IEEE-CIS, Sparkov and a Kaggle e-commerce dataset, then adding synthetic attributes [1]. None of those public sets contains an analyst note, a call summary or a decision rationale.
The gap matters for three common use cases. Case summarization needs long, messy notes with timestamps and handoffs. Typology classification needs analyst-assigned labels such as account takeover (ATO) or authorized push payment (APP) scam rather than a binary fraud flag. Decision support and evaluation need the final disposition plus the evidence the analyst relied on, so a model can be graded against real outcomes; see outcome-labeled evaluation data for how that grading works.
If you only need labeled transactions, start with financial transaction data for AI training. This page covers the narrative case layer and how to buy it safely.
What a usable fraud case record contains
A usable record joins one case to its alert source, its event timeline, its typology and its outcome, with every person replaced by a stable pseudonym. Case management platforms (NICE Actimize, SAS Fraud Management, Featurespace, or in-house tools on Salesforce or ServiceNow) store these as a case header, an activity log and free-text note fields. Ask suppliers to export all three, not just the header.
Core fields to request:
- Alert source: rule ID or model score band, channel (card-not-present, ACH, wire, Zelle-type real-time payment, P2P wallet), and whether the alert was customer-reported or system-generated.
- Typology: account takeover, APP or impersonation scam, first-party fraud (including friendly fraud and bust-out), synthetic identity, mule account, and card-present counterfeit, with the institution's own code list.
- Analyst actions: ordered events such as step-up authentication, card block, device review, beneficiary-bank recall request, and outbound call, each with timestamp and actor role.
- Customer contact notes: call or chat summaries, including what the customer said they were told by the scammer.
- Decision: confirmed fraud, not fraud, customer-authorized, referred, or escalated, plus reimbursement decision where relevant.
- Loss and recovery: gross exposure, amount recovered, and recovery route, stored as amounts or bands.
Features for card and account fraud depend on per-account transaction patterns, so isolated rows lose most of their signal [2]. Case records should therefore carry a pseudonymous account key that joins to a transaction sequence covering a window before and after the alert. Without that join, a copilot cannot learn why an analyst found a payment unusual.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"case_id": "C-7F3A91",
"account_pseudo_id": "ACCT-55e2c1",
"opened_at": "2026-03-04T14:22:00Z",
"alert": {"source": "rule", "rule_id": "RTP_NEW_PAYEE_VELOCITY", "channel": "real_time_payment"},
"typology": {"primary": "app_scam_impersonation", "secondary": null, "code_list_version": "2026.1"},
"events": [
{"t": "2026-03-04T14:25:10Z", "actor": "analyst_L1", "action": "outbound_call", "note": "[PERSON_1] says caller claimed to be bank security; asked to move funds to 'safe account'."},
{"t": "2026-03-04T14:41:02Z", "actor": "analyst_L1", "action": "recall_request_beneficiary_bank"},
{"t": "2026-03-05T09:12:44Z", "actor": "analyst_L2", "action": "decision", "note": "Customer authorized under deception; reimburse per policy."}
],
"decision": "confirmed_scam_customer_authorized",
"loss_band": "1k-5k",
"recovered_band": "0",
"redaction": {"method": "ner_plus_regex_plus_manual_review", "sensitive_categories_removed": ["health", "family"]},
"linked_txn_window_days": [-90, 14]
}
Legal exclusions: SARs, scam-reimbursement rules and consumer-dispute overlap
The firmest exclusion is anything that would reveal whether a Suspicious Activity Report was filed. The Bank Secrecy Act (31 U.S.C. 5318(g)(2)) and FinCEN's bank rule at 31 CFR 1020.320 treat a SAR, and information that would reveal one, as confidential. The underlying transactions and facts are generally treated differently from SAR-revealing information, but where that line falls in a given note is a question for the supplier's BSA officer and counsel, not the buyer.
For a buyer, that translates into concrete field rules. Ask suppliers to drop SAR decision fields, BSA/AML referral flags, SAR narrative drafts, FinCEN filing IDs and any note text that mentions a filing or a "SAR/no-SAR" determination. Fraud cases that escalated to AML should be either excluded or cut at the point of escalation, and the supplier's counsel should sign off on the rule set.
Scam-reimbursement regimes shape labels for buyers outside the US. The UK Payment Systems Regulator's mandatory reimbursement regime for Faster Payments APP scams has applied since October 2024, with a per-claim cap and defined exceptions such as gross negligence; check the current rules as of October 2026 before encoding them. The US has no equivalent mandatory rule for customer-authorized scam payments (Regulation E liability limits address unauthorized transfers), so "reimbursed" on a US scam case usually reflects institution or network policy; map it explicitly before mixing it with UK labels.
Keep consumer-dispute work separate. Reg E and Reg Z error-resolution cases follow regulatory timelines and include non-fraud billing errors, which belong in Reg E and Reg Z dispute investigation records. Merchant-side disputes belong in chargeback representment case data.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Redacting victim narratives without destroying the signal
Victim narratives need more than standard PII removal because they often contain health, family, immigration and relationship details that explain why a person was targeted. Romance and elder-fraud notes are the worst cases. Ask suppliers to remove names, phone numbers, emails, account and card numbers, and also to suppress or generalize sensitive categories such as diagnoses and family members' names.
Automated tools are a first pass, not a guarantee. Microsoft's Presidio project, a common open-source de-identification SDK, states that because it uses trained models there is no guarantee it will find all sensitive information [3]. Fraud notes add specific failure modes: IBANs and routing numbers pasted mid-sentence, scammer phone numbers that are evidence rather than customer data, and analyst shorthand such as "cx said wife has dementia."
Replace rather than delete where possible. Typed placeholders such as [PERSON_1] and [MERCHANT_2], consistent within a case, preserve the narrative structure a summarization model needs. Request the redaction method, the entity list and a sample review result for each delivery.
Checking label quality in analyst decisions
Analyst decisions are noisy labels, and you should measure that noise before training or evaluating on them. Label errors are common even in widely used benchmark test sets and can change which model appears best [4], and fraud dispositions are more subjective than most benchmarks. A first-party claim marked "not fraud" by one analyst may be "customer-authorized" for another.
Practical checks:
- Code-list drift: typology codes are often renamed or split, for example "scam" splitting into impersonation, purchase and investment scam. Ask for the code-list history and effective dates.
- Reversals: cases reopened after a customer complaint or chargeback outcome. Keep the final decision and the original, with timestamps.
- Second-review agreement: where the supplier has QA sampling, request reviewer agreement rates by typology.
- Template text: canned note macros ("Advised customer of scam trends") inflate apparent fluency; see templates and boilerplate in business records.
Treat each case as a trajectory of observations and actions, which also makes it useful for agent evaluation; see reconstructing agent trajectories from ticket and case histories and decision records with rationale.
Buyer request checklist for fraud case data
A precise request names the typologies, channels, period, linkage and exclusions up front, which lets suppliers judge fit before any data moves.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Request element | What to specify | Why it matters |
|---|---|---|
| Typologies | ATO, APP scam, first-party, synthetic identity, mule; target mix | Prevents a set dominated by card-not-present disputes |
| Channels | Card, ACH, wire, real-time payments, P2P | Analyst actions differ by rail (recall vs. chargeback) |
| Period | 24-36 months with code-list versions | Captures typology shifts and taxonomy changes |
| Record scope | Header, full event log, free-text notes, contact summaries | Copilots need the narrative, not just dispositions |
| Linkage | Pseudonymous account key plus transaction window | Restores per-account pattern features [2] |
| Exclusions | SAR fields, AML referrals, law-enforcement requests, Reg E/Z error cases | SAR confidentiality and scope control |
| Redaction | Typed placeholders, sensitive-category suppression, sample review | Automated PII detection can miss entities in free text [3] |
| Documentation | Data card covering sources, labeling process, known gaps | Supports model risk review [5] |
| Delivery | Parquet over an access-controlled share, e.g., Delta Sharing | Shares tables through a controlled protocol rather than loose file copies [6] |
Document everything in a data card: upstream systems, annotation process, intended uses and decisions that affect model performance [5]. Model risk and privacy reviewers will ask for exactly those facts.
How SourceX sources fraud investigation case data
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing and ongoing purchases. Nothing is held in stock, a category is not inventory, and a request does not guarantee a match. Buyers describe the data they need, such as typologies, channels and fields, and SourceX looks for US businesses that hold it; every release is approved by the supplying company.
The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, and delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval.
SourceX does not source scraped web content, standalone contact lists, or generic CCTV or photos, does not train models and does not publish prices. Teams in banking and payments can compare related needs on buyers in finance and buyers in fintech software, or describe your fraud case data requirement. For neighboring categories, browse the industry-specific operational data guide, SOC alert triage decisions and merchant onboarding and KYB decisions, or the AI data hub.
Request fraud case notes and analyst decision data
SourceX serves AI teams wherever they are based and sources investigation data from US companies on request, with every release approved by the supplier and delivered under a license. Describe the typologies, channels, fields and exclusions you need, and SourceX will look for businesses that hold them. Start a fraud case data request.
Sources
- Information Technologies and Mathematical Modelling journal (NMetAU), "Article 2468: building a specialized e-commerce fraud detection dataset" (2026). https://journals.nmetau.edu.ua/index.php/itmm/en/article/view/2468
- arXiv, "A data mining approach using transaction patterns for card fraud detection" (2013). https://arxiv.org/pdf/1306.5547
- Microsoft presidio project (pkg.go.dev), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/pdf/2103.14749
- Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- Databricks, "Introducing Delta Sharing: An Open Protocol for Secure Data Sharing" (2021). https://www.databricks.com/blog/2021/05/26/introducing-delta-sharing-an-open-protocol-for-secure-data-sharing.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.