Industry-specific operational data
AML alert and investigation data for AI: what banks can and cannot license
Quick answer
AML alert data for machine learning splits into three tiers. SAR narratives, SAR filing decisions and any label that reveals a SAR exists cannot leave a bank, because 31 CFR 1020.320(e) makes that information confidential [1]. Alert-level features, scenario metadata and dispositions with no SAR linkage can sometimes be licensed after bank counsel review and de-identification. Everything else, including narrative drafting models, is usually better built inside the bank's environment or on synthetic transaction data.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why SAR confidentiality draws the hard line
The SAR rule is the binding constraint: a bank and its directors, officers, employees and agents may not disclose a SAR or any information that would reveal its existence, outside narrow exceptions such as disclosure to FinCEN, law enforcement and supervisors [1]. The rule sits on top of the Bank Secrecy Act's statutory confidentiality provision, and an AI vendor is not on the list of permitted recipients. As of October 2026, FinCEN has posted a joint statement on SAR confidentiality (file dated September 2026) [2]; read it before scoping any project.
The practical consequence is broader than "no SAR forms." Fields such as sar_filed, sar_decision, escalated_to_sar_committee, a case status of "SAR filed," a BSA ID, or a narrative that quotes the filing all reveal existence. So does a label set where the positive class is defined as "cases that became SARs," even if every narrative is stripped out.
The regulation's rules of construction generally treat the underlying facts, transactions and documents behind a SAR differently from the SAR itself [1]. That does not make a SAR-derived dataset licensable. A dataset assembled by selecting transactions because they were reported still discloses the filing through its selection logic.
Why 314(b) is not a channel to AI vendors
Section 314(b) lets eligible financial institutions and associations of them share information with each other about possible money laundering or terrorist activity, with a safe harbor from liability under 31 CFR 1010.540 [7]. As of October 2026, the safe harbor depends on registration with FinCEN, verifying the counterparty is registered, using the information for AML purposes and safeguarding it; confirm the current text. A model developer that is not itself an eligible financial institution with an AML program cannot participate, and training a commercial model is not the purpose the rule protects.
Consortium-style AML products therefore structure themselves differently: they run inside participating banks, exchange model parameters or risk scores rather than raw case files, or rely on bank-signed data agreements reviewed for SAR and privacy exposure. Treat any offer of "314(b) data" for model training as a red flag.
What can leave a bank after review
Material can be licensable when it describes the monitoring process rather than the reporting outcome. Even then, customer data carries Regulation P limits: a recipient of nonpublic personal information under an exception may use it only for the purpose it was received [3], which is why de-identification and a clear license purpose matter. The bank's third-party risk program will also apply to you as a recipient [4].
Illustrative example: invented to show structure; it does not describe an available dataset.
| Data element | Typical source system | Licensing posture | Main risk |
|---|---|---|---|
| SAR narrative, SAR form fields, BSA ID | SAR e-filing records, case management | Cannot leave the bank | 31 CFR 1020.320(e) [1] |
| Label "SAR filed / not filed", SAR committee minutes | Case management workflow | Cannot leave the bank | Reveals existence of SAR |
| Investigation case notes and RFI responses | Case management (e.g., Actimize, Oracle FCCM, Verafin, custom) | In-bank use only, in most programs | Notes often reference filing decisions |
| Alert-level features: scenario ID, threshold, score, amounts bucketed, segment | Transaction monitoring engine | Possible after counsel review and de-identification | Re-identification via rare amounts and dates |
| Alert disposition at L1 (closed, escalated) with no case linkage | Alert queue | Possible with care | Escalation can proxy for SAR if filings are rare |
| Scenario and rule library metadata, tuning history | Model governance inventory | Often licensable | Reveals detection thresholds; security review |
| Investigator workflow timings, queue states | Workflow logs | Often licensable | Staffing data, employee privacy |
| Sanctions screening hits | Screening system | Separate regime; out of scope here | OFAC handling rules |
Two adjacent categories have their own pages: fraud operations notes in fraud investigation case notes and analyst decisions and onboarding files in KYC and CDD case review files. Sanctions screening is a separate subject and not covered here.
Why disposition labels mislead triage models
A disposition of "closed, no further action" is not a ground-truth label for benign activity. It records that an investigator, working against a threshold, a queue target and a backlog, decided the alert did not warrant escalation. Dispositions shift when scenarios are retuned, when staffing changes, or after a regulatory finding drives a lookback.
Researchers make the same point from the other side: fraud-detection studies build training sets from open and synthetic sources because real labeled data is scarce [5], and in real AML data many laundering transactions are never detected, so labels are incomplete by construction. A false-positive reduction model trained on raw dispositions learns the bank's historic triage policy, including its blind spots.
Practical controls before training:
- Stratify by scenario version and threshold epoch; never pool alerts across a tuning change without a version feature.
- Record queue age at disposition; alerts closed in bulk near quarter end carry weaker labels.
- Separate L1 auto-closures and rule suppressions from human dispositions.
- Hold out a period after any lookback or consent order as an evaluation slice.
- Measure coverage of rare typologies such as structuring, funnel accounts and trade-based patterns; see long-tail and edge-case coverage.
How to build SAR narrative and copilot models without exporting SARs
Narrative drafting and investigation copilots are trained where the data already lives: inside the bank's environment. The common patterns are a vendor model fine-tuned on bank infrastructure, with weights that stay in the bank's tenancy; retrieval over the bank's own case archive at inference time; and evaluation sets built and scored by bank staff. The vendor supplies the base model, training code and evaluation harness, not the data pipeline.
Differentially private fine-tuning and DP synthetic text can reduce memorization risk when a model must learn from sensitive narratives [6]. Treat them as a control inside the bank, not as a license to ship narratives out. A model trained on SARs that can regurgitate a filing remains a disclosure risk.
For bank-side governance, the model will fall under the institution's model risk program; see model risk management for third-party training data in banking.
Where synthetic AML transaction data fits
Synthetic transaction sets are the default for pretraining and benchmarking graph and tabular AML models outside a bank. The IBM Research and ETH Zurich datasets released with the NeurIPS 2023 paper (often called AMLworld) come from an agent-based generator calibrated to resemble real transactions, with complete ground-truth laundering labels, published in the NeurIPS 2023 Datasets and Benchmarks Track (Altman et al.) [8]. They are useful for architecture selection and pattern detection, but they have no investigator notes, no alert queues and no disposition behavior.
A workable split is synthetic data for detection pretraining, licensed process data (scenario metadata, de-identified alert features, workflow timings) for triage and operations modeling, and in-bank fine-tuning for anything touching cases or narratives.
A request template for AML operations data
The cleanest request names fields, exclusions and the training location up front, so bank counsel can review it quickly.
Illustrative example: invented to show structure; it does not describe an available dataset.
use_case: alert triage prioritization model (L1 queue)
records: transaction monitoring alerts, 24 months
fields_requested:
- alert_id (re-keyed token)
- scenario_id, scenario_version, threshold_value
- alert_score, amount_bucket, txn_count_bucket, customer_segment
- l1_disposition (closed | escalated), disposition_timestamp_week
- queue_age_days, investigator_role (no names)
explicit_exclusions:
- any SAR, SAR decision, SAR committee or BSA ID field
- case notes, RFI text, narratives
- alerts selected because of a filing outcome
- sanctions screening hits
deidentification: customer and account IDs tokenized; dates coarsened to week
training_location: buyer environment (process data only)
bank_signoff: BSA officer and legal review recorded before release
How SourceX handles AML-adjacent requests
SourceX sources operational datasets from US companies on request and manages licensing and ongoing purchases; categories are not inventory, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect.
Every release is approved by the supplying company, and nothing is contracted until a supplier agrees. For context on the wider category, see financial services records, financial transaction data and the finance buyer page; compliance consultancies can start at compliance consulting buyers. Buyers can describe the AML process data they need at any stage.
Sourcing AML alert data for machine learning
SourceX looks for US businesses that hold the operational data you describe and runs each request through Find, Assess, Agree, Transact and Manage, with rights review before any license. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Explore more in the industry data hub, then submit your request to SourceX.
Frequently asked questions
Can a bank share SAR data with an AI vendor for training?
No. The SAR confidentiality rule bars disclosure of a SAR or information revealing its existence outside listed exceptions, and AI vendors are not among them [1]. Train on SAR material only inside the bank, under its own controls.
Is a de-identified SAR narrative still a SAR?
Generally yes for this purpose. Removing names does not change that the text is the content of a filing, and its release would reveal that a filing exists [1].
Are AML alert dispositions useful training labels?
They are useful for modeling triage behavior, not for ground truth on laundering. Version them by scenario and threshold, and evaluate against held-out periods.
Sources
- Electronic Code of Federal Regulations (eCFR), "31 CFR 1020.320 - Reports by banks of suspicious transactions". https://www.ecfr.gov/current/title-31/subtitle-B/chapter-X/part-1020/subpart-C/section-1020.320
- Financial Crimes Enforcement Network (FinCEN), "Joint Statement on SAR Confidentiality" (2026). https://www.fincen.gov/system/files/2026-09/Joint-Statement-on-SAR-Confidentiality.pdf
- Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
- Board of Governors of the Federal Reserve System, "Interagency Guidance on Third-Party Relationships: Risk Management" (2023). https://www.federalreserve.gov/frrs/guidance/interagency-guidance-on-third-party-relationships.htm
- Information Technology and Mathematical Modelling (journal, NMetAU), "Fraud detection research merging public and synthetic datasets". https://journals.nmetau.edu.ua/index.php/itmm/en/article/view/2468
- arXiv, "Differentially Private Knowledge Distillation via Synthetic Text Generation" (2024). https://arxiv.org/html/2403.00932v2
- Electronic Code of Federal Regulations (eCFR), "31 CFR 1010.540 - Voluntary sharing of information among financial institutions". https://www.ecfr.gov/current/title-31/subtitle-B/chapter-X/part-1010/subpart-E/section-1010.540
- arXiv, "Realistic Synthetic Financial Transactions for Anti-Money Laundering Models" (2023). https://arxiv.org/abs/2306.16424
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.