Privacy, de-identification and sensitive data
Financial records in AI training data: GLBA, nonpublic personal information and de-identification
Quick answer
GLBA does not ban training on financial records, but it follows the data. If a bank, lender or insurer shares nonpublic personal information (NPI) with you, Regulation P's reuse and redisclosure limits bind you as the recipient, and data received under a processing or service-provider exception generally cannot be repurposed for model training [1][2]. Information that does not identify a consumer, such as aggregate or blind data without identifiers, sits outside NPI, so a well-evidenced de-identification step is usually the cleanest path [3].
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
GLBA binds the recipient, not only the bank
The reuse and redisclosure rule attaches to whoever receives NPI from a nonaffiliated financial institution, whether or not that recipient is itself a financial institution [1][2]. The FTC's compliance guide states this plainly: a business with no consumers or customers of its own can still be bound by the limits on information it receives [2]. Regulation P (12 CFR Part 1016) is the CFPB's version and covers most banks and nonbank financial companies; the FTC's 16 CFR Part 313 now reaches mainly certain motor vehicle dealers, and the SEC's Regulation S-P (17 CFR Part 248, Subpart A) covers broker-dealers, investment companies and registered advisers, so which one governs depends on the originating entity [4].
For an AI developer, that means the question is never just "are we a financial institution?" It is "how did each record leave the institution, and what did that pathway allow?" A fintech, a model vendor and a data intermediary can all inherit the same constraint.
How the data left the institution decides what you can do with it
The pathway the NPI took out of the originating institution sets the ceiling on your reuse, and two very different rules apply [1][2]. Data shared under the processing and servicing exceptions (12 CFR 1016.14) or the other listed exceptions (1016.15) may be used and disclosed by the recipient only in the ordinary course of business to carry out the purpose for which it was received [1]. Training a general-purpose or cross-client model is rarely that purpose.
Data shared outside an exception, after the institution gave its privacy notice and the consumer did not opt out under 1016.10, is different. There the recipient "steps into the shoes" of the institution and may disclose to the same kinds of parties the institution itself could [2]. Data shared with a service provider under 1016.13 travels with a contract that limits the provider's use to the services it performs [3].
Illustrative example: invented to show structure; it does not describe an available dataset.
| How the records reached the supplier | What Regulation P lets the recipient do | Fit for licensing into AI training |
|---|---|---|
| Processing or servicing exception (1016.14), e.g. a card processor handling authorizations | Use and disclose only to carry out the purpose received [1] | Poor as identifiable NPI; requires de-identification first |
| Other exceptions (1016.15), e.g. fraud prevention, legal process | Same ordinary-course limit [1] | Poor; fraud-model reuse needs careful scoping by counsel |
| Service provider under 1016.13 (vendor contract) | Contractual limit to the services performed [3] | Poor unless the institution's contract expressly permits it |
| Disclosed after notice and no opt-out (1016.10) | Steps into the shoes of the institution [2] | Possible, but tied to the institution's notice and opt-out records |
| The supplier is the originating institution (its own customers) | Bound by its own notice and opt-out obligations [4] | Possible; check what the privacy notice says about sharing |
| Aggregate or blind data with no identifiers | Outside NPI by definition [3] | Strongest pathway, if the de-identification holds up |
The most common failure mode is a supplier that is itself a recipient, such as a payments processor or loan servicer, offering "its" transaction data for licensing when the data arrived under 1016.14. That supplier cannot pass on rights it never had.
When de-identified financial data falls outside NPI
Regulation P defines NPI through "personally identifiable financial information" and excludes information that does not identify a consumer, such as aggregate information or blind data that lacks personal identifiers like account numbers, names or addresses [3]. Unlike HIPAA, GLBA has no Safe Harbor list or Expert Determination route, so the regulation gives you a concept, not a method. Counsel has to judge whether a specific transaction-level file actually "does not identify" anyone.
Three patterns create disputes in practice. First, a persistent pseudonymous key, such as an HMAC of the primary account number, keeps a longitudinal customer history intact; if the supplier retains the key or a lookup table, many reviewers treat the file as still identifiable. Second, merchant descriptors, ACH originator names, wire memo fields and free-text dispute notes routinely carry names and account fragments that column-level masking misses. Third, fine-grained timestamps, ZIP+4 and exact amounts let a holder of auxiliary data link records back to people, which is the mosaic problem covered in combining de-identified datasets and linkage risk.
For methods, NIST SP 800-188 is the most useful public reference on de-identification techniques and governance, including risk assessment and release models [7]. Our step-by-step guide to de-identifying financial transaction data covers field-level treatment of PANs, routing numbers, descriptors and amounts.
State privacy laws: GLBA exemptions are not uniform
Many state comprehensive privacy laws carve out GLBA-covered data, but the carve-outs differ in scope and the state de-identification standards differ too [5]. Some states exempt at the entity level, so a GLBA-regulated financial institution sits outside the statute entirely. Others, notably California, exempt at the data level: only information collected under GLBA is excluded, and the same company's other data stays in scope [5].
That matters for AI buyers because a dataset can drift out of the GLBA exemption once it leaves the regulated pipeline or is merged with non-GLBA sources, such as app telemetry or marketing data. California's definition of "deidentified" then becomes relevant: reasonable technical measures, a public commitment not to re-identify, and contractual obligations on every recipient [6]. Drafting to that standard is a common conservative choice when the data's state-law status is unclear.
What to request from a supplier of financial records
Ask for evidence of the data's GLBA pathway and the de-identification method before you negotiate price or scope. The list below is a practical working checklist for counsel and data teams.
Illustrative example: invented to show structure; it does not describe an available dataset.
Financial records diligence checklist
- Originating entity and regulator. Which institution collected the data, and whether Regulation P, 16 CFR 313 or Regulation S-P applies [4].
- Pathway statement. Whether the supplier holds the data as the originating institution, under 1016.13, 1016.14, 1016.15 or after notice and opt-out [1][3].
- Privacy notice and opt-out records. The notice versions in force during the collection window and how opt-outs were honored, if identifiable data is in scope [2].
- Upstream contracts. Processor, servicer or core-banking vendor agreements that restrict secondary use.
- De-identification specification. Field-by-field treatment (PAN, routing and account numbers, SSN or TIN, names in descriptors, addresses, device IDs), key custody for any tokens, and date and amount generalization [7].
- Free-text handling. How dispute narratives, call notes, collections notes and memo fields were scanned and redacted, with a sampled error rate.
- Re-identification risk assessment. Method, assumed attacker and auxiliary data, and residual risk; see re-identification risk assessment for licensed datasets.
- State-law position. Which state exemptions the supplier relies on and whether California-style deidentified conditions are met [5][6].
- Re-identification prohibition. The clause you will be asked to sign; compare against re-identification prohibition clauses in data licenses.
For the full document set, see the de-identification evidence package checklist.
Fraud, dispute and agent use cases need narrower scoping
Fraud and dispute models are where GLBA reasoning gets subtle, because the 1016.15 exceptions include fraud prevention but tie use to the purpose for which data was received [1]. A model trained for one institution's fraud program is a different purpose from a commercial fraud model sold to many clients. Counsel should document which purpose the license supports.
Agent and SFT datasets built from servicing transcripts, collections notes or chargeback case files carry the highest free-text risk. Labels drawn from historical credit or claims decisions also inherit bias, covered in auditing historical decision bias in operational labels. Banks deploying the model will also apply model risk management to your training data, as described in model risk management for third-party training data.
Commitments you make downstream still count
Once financial records are in your pipeline, your own representations become enforceable obligations. FTC staff have warned AI companies that using customer data for undisclosed purposes, including training, can violate laws the FTC enforces [8]. If you resell or host models for banks, keep your contract promises about client data separate from any licensed training corpus, and track lineage so you can show which records trained which model version.
GLBA also expects security for customer information under the Safeguards Rule (16 CFR Part 314), so keep any identifiable residue under access controls and logging [9], and treat de-identified files as sensitive until the risk assessment is accepted. Examiners review reuse and redisclosure as part of GLBA privacy exams, so originating institutions will want evidence that your use stays within the permitted pathway [3].
How SourceX handles financial records for AI buyers
SourceX sources operational datasets, including finance workflows, support histories and documents, from US companies on request, and manages the licensing process; categories are not inventory and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe the financial records you need without naming businesses.
Request de-identified financial records for AI training
Describe the bank, lending, payments or insurance records your model needs, and SourceX will look for US businesses that hold them, assess data and licensing permissions, and agree allowed uses in a license before anything is delivered. Every release is approved by the supplying company, and nothing is contracted until a supplier agrees. Start a buyer request.
Related: privacy and de-identification hub, GLBA and AI data licensing, license financial transaction data, all AI data guides.
Sources
- Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
- Federal Trade Commission, "How To Comply with the Privacy of Consumer Financial Information Rule of the Gramm-Leach-Bliley Act". https://www.ftc.gov/business-guidance/resources/how-comply-privacy-consumer-financial-information-rule-gramm-leach-bliley-act
- Consumer Financial Protection Bureau, "CFPB Laws and Regulations: GLBA Privacy" (2016). https://files.consumerfinance.gov/f/documents/102016_cfpb_GLBAExamManualUpdate.pdf
- Consumer Financial Protection Bureau, "12 CFR 1016.1 - Purpose and scope". https://www.consumerfinance.gov/rules-policy/regulations/1016/1/
- National Law Review, "Finding the Delta: Understanding the Differences in State Deidentification Standards". https://www.natlawreview.com/article/finding-delta-understanding-differences-state-deidentification-standards
- California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- National Institute of Standards and Technology, "NIST SP 800-188: De-Identifying Government Datasets: Techniques and Governance" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
- Federal Trade Commission / eCFR, "Standards for Safeguarding Customer Information (16 CFR Part 314)". https://www.ecfr.gov/current/title-16/chapter-I/subchapter-C/part-314
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.