Skip to content

Industry-specific operational data

Reg E and Reg Z dispute investigation records for dispute-intake and investigation AI

Quick answer

Bank dispute case data for AI means issuer-side error-resolution files: the customer's intake narrative, the claim type, transaction metadata, the provisional credit decision, investigator notes and evidence, the final determination and the letters sent. Debit and ACH claims run on Regulation E clocks and credit card billing errors on Regulation Z clocks. Each case is a timestamped trajectory, so buy whole cases with event dates and path-specific outcome labels, licensed from the bank or processor that ran the investigations.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What a Reg E or Reg Z dispute case file contains

A usable dispute record is the full case, from first contact to closing letter, not a row in a transaction table. Fraud-labeled transaction tables answer "was this transaction bad"; dispute files answer "what did the customer claim, what did the bank do, by when, and why." That second question is what intake agents, investigation copilots and letter-drafting models learn from. For the transaction-level feeds themselves, see licensing financial transaction data for AI training.

Most issuers hold this data across several systems: a case management tool (often a dispute module in the core or card platform), the card processor's dispute queue, the contact center platform holding call transcripts or chat logs, and a document store for affidavits, receipts and merchant correspondence. Expect to join records across these systems on a case key. The guide on packaging linked records from multiple business systems covers how to ask for that join.

The core components to request:

  • Intake: channel (phone, chat, branch, online form), verbatim or summarized narrative, the agent's claim-type selection, date of first notice and whether written confirmation was requested.
  • Transaction context: amount, posting and transaction dates, merchant category code, merchant descriptor, card-present or card-not-present, transfer type (POS, ATM, ACH, P2P) and account age.
  • Investigation: step log with timestamps, evidence received, merchant or network contacts, analyst notes and any reopen events.
  • Decisions: provisional credit granted, amount and date; final determination; reversal of provisional credit; customer notices and their send dates.

How Regulation E and Regulation Z timelines shape the labels

The regulatory clock is the backbone of the label schema, because a model that drafts letters or prioritizes queues must know which deadline applies. Under Regulation E, an institution generally has 10 business days after receiving a notice of error to investigate, and may take up to 45 days only if it provisionally credits the account within those 10 business days [1]. It must tell the consumer the amount and date of the provisional credit within two business days and report results within three business days of finishing [1].

The rule also allows withholding up to $50 of provisional credit in some unauthorized-transfer cases, and sets longer windows: up to 90 days to investigate for new-account, point-of-sale debit card and foreign-initiated transfers, plus 20 business days instead of 10 for new accounts [1]. When an error is found, correction includes crediting interest and refunding institution-imposed fees where applicable [2]. Each of these branches should be visible in the data as fields, not buried in notes.

Regulation Z runs a different clock for credit card billing errors. The consumer's written notice must arrive within 60 days after the creditor sent the first periodic statement reflecting the error, the creditor acknowledges within 30 days, and resolution must come within two complete billing cycles and no later than 90 days [3]. While the dispute is pending, the creditor faces limits on collecting the disputed amount and reporting it as delinquent [3]. As of October 2026, confirm the current text of both sections on the CFPB and eCFR pages before encoding deadlines into training targets.

Practical consequence: store every clock-relevant timestamp (notice received, written confirmation received, provisional credit posted, notice sent, determination, results letter sent) and the applicable regime. Without them you cannot train timeline-aware agents or measure compliance in evaluation.

Why outcome labels must separate resolution paths

A single "won/lost" outcome field is the most common defect in dispute data. A claim can be resolved internally (error found and corrected), passed to the card network as a chargeback, recovered from the merchant directly, denied after investigation, or withdrawn by the customer. A chargeback that later loses at representment can still end with the issuer absorbing the loss and keeping the customer whole.

Model each path separately: resolution_path, customer_outcome and issuer_financial_outcome. Intake classifiers and outcome predictors trained on a collapsed label learn that "merchant dispute" means "customer loses," which is often wrong. The network leg of the same dispute is a separate dataset with reason codes and compelling evidence; see chargeback representment case data.

Reopened and reversed cases deserve their own flag. They carry the error-recovery signal that copilots need, as described in rework, reversals and reopened cases.

Illustrative case record schema

The schema below shows the minimum structure that supports intake classification, investigation copilots, letter drafting and outcome prediction from the same file.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldTypeExample valueTraining use
case_idstring (salted hash)c_7f3a91Join key across systems
regimeenumREG_E / REG_ZSelects deadline logic
intake_channelenumphoneChannel-specific intake agents
intake_narrativetext (redacted)"I didn't make the [MERCHANT] charge on [DATE]"Claim-type classification, SFT
claim_typeenumunauthorized / wrong_amount / goods_not_received / duplicate / atm_cash_not_dispensedClassifier target
txn_typeenumpos_debitWindow selection
account_age_daysinteger22New-account branch
notice_received_attimestamp2026-03-02T14:10ZClock start
provisional_creditstruct{granted: true, amount: 184.20, posted_at: ..., withheld: 50.00}Provisional credit decision data
investigation_stepsarray of events[{step: "merchant_contact", at: ...}]Copilot trajectories
evidence_refsarray["receipt_img_01"]Multimodal evidence review
determinationenumerror_found / no_error / partialOutcome target
resolution_pathenuminternal / network_chargeback / merchant_refund / withdrawnPath-aware labels
lettersarray of docs (redacted)results notice, reversal noticeLetter drafting
reopenedboolean + reasontrue, "new evidence"Error recovery

Document each field and enum in a data dictionary using a template like the one for licensed dataset deliveries.

De-identification and GLBA reuse limits for dispute data

Dispute files are dense with nonpublic personal information, so de-identification and reuse permissions decide whether a deal is viable. Under the GLBA Privacy Rule, a business that receives nonpublic personal information under an exception may use and disclose it only in the ordinary course of business to carry out the purpose for which it was received [5]. The FTC also explains that a recipient "steps into the shoes" of the originating institution for redisclosure, and that account numbers may not be shared for marketing, with encrypted numbers allowed only where the recipient cannot decode them [4].

For AI training data that translates into concrete rules. Remove primary account numbers entirely rather than masking to last four, consistent with PCI DSS practice, and replace account and case identifiers with irreversible salted hashes that the supplier keeps. Redact names, addresses, phone numbers and email addresses from narratives, transcripts and letters, and replace them with typed placeholders so the language model still learns structure.

Free text is the hard part. Narratives contain merchant names tied to small towns, unusual amounts and dates that, combined, can single out a customer; re-identification of supposedly de-identified data is well documented [6]. Ask the supplier to describe the method, the residual-risk review and the sample check, and read the guide on how to de-identify financial transaction data before signing.

Coverage and quality checks before you license

Check coverage against the decisions your model must make, not against total case counts. Dispute volume is dominated by a few claim types, so rare branches (foreign-initiated transfers, new-account extensions, partial credits, ATM cash-not-dispensed, reopened cases) can be thin. The page on long-tail and edge-case coverage explains how to set minimums per slice.

Illustrative example: invented to show structure; it does not describe an available dataset.

Buyer request checklist for dispute investigation records

  1. Regimes: Reg E deposit and debit, Reg Z credit card, or both, flagged per case.
  2. Claim-type taxonomy as used by the issuer, with its mapping history if it changed.
  3. All clock-relevant timestamps, in UTC, with business-day calendar used.
  4. Provisional credit fields including partial withholding and reversal.
  5. Separate resolution-path, customer-outcome and issuer-financial-outcome labels.
  6. Investigation step log, not only final notes.
  7. Letters and notices as text, with template IDs.
  8. Reopen and complaint linkage flags.
  9. De-identification method, residual-risk review and sample-check results.
  10. Written confirmation of permitted AI training use and redisclosure basis.

Watch for known failure modes: claim types relabeled by investigators after intake (keep both values), letter text generated from templates with no free-text variation, and missing cases that escalated into complaints. Complaints that cite a dispute are a distinct dataset covered in bank complaint records for AI, and fraud analyst work lives in fraud investigation case notes.

Where dispute records fit among industry datasets

Dispute investigation records sit between transaction data and case-management histories. They belong to the broader family of industry-specific operational data for AI, and they are a natural input for agent trajectory work alongside enterprise workflow datasets. Teams evaluating several banking data types can start from the finance buyer overview.

SourceX sources operational datasets from US companies on request, including support histories and finance workflows, and manages the licensing and ongoing purchases. Buyers describe the data they need, not the businesses, and every release is approved by the supplying company; you can describe dispute case requirements to SourceX.

Licensing Reg E and Reg Z dispute case data

SourceX looks for US businesses that hold the dispute investigation records you describe, rights-reviews each dataset for ownership and consents, and delivers under a license defining records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Data is sourced on request and a request does not guarantee a match; start at SourceX for buyers.

Frequently asked questions

Can I train on dispute data that includes card network chargeback outcomes?

Yes, if the license covers those fields, but keep the network leg separate from the issuer's Reg E or Reg Z determination. Network rules and reason codes differ from the consumer-protection timelines, and mixing them corrupts both labels.

Should intake narratives be summarized or verbatim?

Verbatim, redacted text is more valuable for intake agents because it preserves how customers actually describe problems. Summaries written by agents carry the agent's claim-type bias. Ask for both where available.

Is a sample of dispute cases enough to validate a supplier?

A sample shows field completeness and redaction quality, but not slice coverage. Ask for a field-level profile across the full candidate set, including claim-type and regime distributions, alongside any sample.

Sources

  1. Electronic Code of Federal Regulations (eCFR), "12 CFR 1005.11 - Procedures for resolving errors (Regulation E)". https://www.ecfr.gov/current/title-12/chapter-X/part-1005/subpart-A/section-1005.11
  2. Consumer Financial Protection Bureau, "Official Interpretation of 1005.11 (Comment 11)". https://www.consumerfinance.gov/rules-policy/regulations/1005/interp-11/
  3. Consumer Financial Protection Bureau, "12 CFR 1026.13 - Billing error resolution (Regulation Z)". https://www.consumerfinance.gov/rules-policy/regulations/1026/13/
  4. Federal Trade Commission, "How To Comply with the Privacy of Consumer Financial Information Rule of the Gramm-Leach-Bliley Act". https://www.ftc.gov/business-guidance/resources/how-comply-privacy-consumer-financial-information-rule-gramm-leach-bliley-act
  5. Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
  6. National Institute of Standards and Technology, "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data