Skip to content

Privacy, de-identification and sensitive data

Substance use disorder records (42 CFR Part 2) in AI training data after the 2024 rule

Quick answer

Part 2 substance use disorder (SUD) records can enter an AI training set in practice only after they are de-identified to the HIPAA standard in 45 CFR 164.514(b), which the February 2024 final rule adopted for Part 2 [1][3]. The rule took effect April 16, 2024, and compliance was due February 16, 2026 [1][6]. Buyers should require proof that Part 2 records were flagged before de-identification, that SUD counseling notes were excluded, and that stricter state laws were checked.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What the 2024 Part 2 rule changed for data that leaves a program

The 2024 rule moved Part 2 much closer to HIPAA without merging the two regimes. HHS, through SAMHSA and the Office for Civil Rights (OCR), announced the rule on February 8, 2024, as the CARES Act section 3221 alignment of Part 2 with HIPAA and HITECH [2]. It was published at 89 FR (No. 33) on February 16, 2024 [1]. As of October 2026 the compliance date has passed, so supplier records disclosed after February 16, 2026 should reflect the new consent, notice and de-identification mechanics [1][6].

The changes that matter for a data buyer fall into four groups. Patients can now sign a single consent for future uses and disclosures for treatment, payment and health care operations, and HIPAA covered entities and business associates that receive records under that consent may redisclose them as HIPAA permits [7]. The definition of de-identified information and the public health disclosure route now reference the HIPAA de-identification standard [1][5]. Penalties and breach notification were aligned with HIPAA and HITECH, with enforcement responsibility moving toward OCR [1][2].

Two protections remain stricter than HIPAA. Part 2 still restricts use of SUD records in civil, criminal, administrative or legislative proceedings against the patient, and the rule created SUD counseling notes as a distinct category that needs its own separate consent [7]. HHS published no AI-specific guidance on Part 2 data in the materials we reviewed, so model training should be analyzed through the de-identification route rather than any training-specific exception.

Why de-identification is the practical path into a training set

De-identification is the route because a TPO consent does not obviously cover licensing records to a third party to train a commercial model. Health care operations under HIPAA covers a covered entity's own quality, training and business-management activities. It does not cover selling or licensing patient-level records to an outside AI developer, and a sale of PHI needs separate authorization. Research disclosures under Part 2 carry their own conditions and are a poor fit for an ongoing commercial license.

Once records meet the HIPAA de-identification standard, the 2024 rule treats them as outside the identifiable-record restrictions, in the same way HIPAA treats de-identified data as no longer individually identifiable health information [1][3]. Practice guides describe the rule as applying the HIPAA standard to Part 2 data shared for public health, analytics and research [5]. That is the narrow door a training dataset has to pass through.

The HIPAA standard offers two methods [4]:

  • Safe Harbor (164.514(b)(2)): remove 18 identifier types for the patient and relatives, employers or household members, and have no actual knowledge that the remainder could identify someone [3][4].
  • Expert Determination (164.514(b)(1)): a qualified expert applies statistical and scientific methods, documents the analysis and concludes the risk of re-identification is very small for the anticipated recipient and context [4].

A limited data set under 164.514(e) is not de-identified and keeps dates and some geography, so it is not a substitute here [3]. For SUD data in particular, see HIPAA Safe Harbor vs Expert Determination for AI training data for when each method suits a buyer's use.

Where Part 2 data hides inside claims, notes and revenue-cycle data

Part 2 records are rarely delivered as a labeled table, which is why segmentation is the central buyer risk. A payer or revenue-cycle management (RCM) export may carry SUD information in several places at once, and a single unflagged column can defeat an otherwise clean de-identification. Part 2 applies to records from federally assisted programs that hold themselves out as providing SUD diagnosis, treatment or referral, plus lawful holders that receive those records [1].

Typical locations in healthcare administration data:

  • Diagnosis codes: ICD-10-CM F10-F19 (disorders due to psychoactive substance use, including F17 nicotine dependence, which most screens treat separately) on 837P/837I claims, encounter records and problem lists.
  • Procedure and revenue codes: HCPCS H-codes for behavioral health and SUD services, opioid treatment program G-codes, and UB-04 revenue codes for detox or residential treatment.
  • Pharmacy data: NDCs for buprenorphine, buprenorphine-naloxone, naltrexone and other recovery medications in NCPDP claims.
  • Provider identifiers: NPIs and taxonomy codes for opioid treatment programs or SUD facilities, which reveal the nature of care even when diagnoses are stripped.
  • Free text: prior authorization narratives, utilization review notes, denial letters, call-center summaries and clinical notes that mention treatment, drug screens or recovery.
  • Remittance and denial data: 835 claim adjustment reason codes and appeal correspondence tied to SUD services.

Free text carries the most residual risk. A note can name a treatment facility, a counselor or a court referral long after structured fields were cleaned. Our guide to de-identifying clinical free text for LLM training covers detection and surrogate replacement in more detail, and the utilization management review records page shows where SUD content often appears in payer workflows.

Segmentation checks to require from a supplier

Require the supplier to show that Part 2 records were identified before de-identification, not inferred afterward. Most failures happen when a supplier de-identifies a mixed dataset with a generic PHI pipeline and never asks whether any record came from a Part 2 program. The rule's notice-to-accompany-disclosure mechanics and consent tracking give a supplier the hooks to flag records at intake, if its systems use them [1][7].

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckWhat to ask forFailure mode it catches
Source flaggingHow records from Part 2 programs or received under Part 2 consent are tagged (data segmentation for privacy labels, consent-registry joins, facility lists)SUD records enter the pipeline untagged
Code-based screenThe ICD-10, HCPCS, NDC and NPI taxonomy lists used to find SUD content, with version datesIndirect SUD signals left in structured fields
Counseling notesConfirmation that SUD counseling notes were excluded, not de-identifiedA category needing separate consent slips in
Free-text scanDetector, recall measured on a labeled sample, and reviewer sign-offFacility names and referral sources in notes
Method recordSafe Harbor checklist or the Expert Determination report, scope and dateUnsupported "de-identified" label
State law reviewWhich state SUD, mental health or HIV laws were checkedA stricter state rule still applies [5]
Linkage limitsWhether tokens or join keys allow linkage to other licensed dataRe-identification by combining datasets
Disclosure timingWhether records were disclosed before or after February 16, 2026Old consents interpreted under new rules [1]

Linkage deserves its own review because SUD status is exactly the kind of attribute a re-identifier wants. NIST's survey of de-identification documents real cases in which supposedly de-identified data was re-identified [8]. If you plan to combine a claims extract with other licensed data, read combining de-identified datasets without re-identifying anyone before the Expert Determination is scoped, since the expert must assess the recipient's actual environment [4].

An illustrative de-identified record and what was removed

A compliant training record keeps clinical and operational signal while dropping or generalizing anything that identifies the patient, the program or the people around them. The example below shows a prior authorization record after an Expert Determination workflow, with the transformations recorded in a sidecar file rather than in the record itself.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "tok_7f3a91",
  "part2_origin_flag": true,
  "deid_method": "expert_determination",
  "deid_report_ref": "ED-2026-014",
  "member_age_band": "35-44",
  "state_region": "Census division 5",
  "service_month_offset": 0,
  "dx_codes": ["F11.20"],
  "hcpcs_codes": ["H0020"],
  "request_type": "continued_stay",
  "decision": "approved_partial",
  "review_note": "[FACILITY] requested 14 additional days. Reviewer approved 7 pending [DATE] reassessment.",
  "excluded_categories": ["sud_counseling_notes", "court_referral_text"]
}

Notice the explicit part2_origin_flag. Keeping the flag lets the buyer's governance team apply stricter retention, access or evaluation rules to those rows, and lets them be removed if the method is later questioned. Whether retaining the diagnosis code is acceptable depends on the expert's analysis of the release context; under Safe Harbor, codes may remain but dates must be reduced to year and geography generalized [3][4].

Contract and governance terms that follow the data

The license should restate the de-identification method and forbid re-identification, because the HIPAA standard depends on the recipient's environment as much as on the data. Expert Determination reports are usually scoped to a named recipient and a defined set of controls, so a license that allows resale or open release can void the analysis [4]. Ask for the report scope in writing.

Clauses buyers commonly negotiate for SUD-derived data include a no-re-identification and no-linkage covenant, limits on combining with consumer data, notice if the supplier learns of a source-side consent problem, and a deletion or suppression path for records later found to be identifiable. Model outputs deserve attention too: a model fine-tuned on prior authorization narratives can memorize rare phrasing, so evaluate extraction risk before deployment. Our HIPAA and AI training data overview and the PHI glossary entry set out the base definitions these terms rely on.

How SourceX handles health records that may include Part 2 data

SourceX sources operational datasets from US companies on request, including support, finance and workflow records, and manages licensing and ongoing purchases. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Health records require HIPAA de-identification by Safe Harbor or Expert Determination, and personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked, though no method is perfect. Diligence materials on source, rights, preparation and allowed use are prepared per dataset, which is where Part 2 segmentation questions belong; describe your requirements on the SourceX buyer intake.

For the wider set of privacy questions, start at the privacy and de-identification hub or the AI data overview.

Request de-identified healthcare operations data

SourceX looks for US businesses that hold the data you describe, and every release is approved by the supplying company; a request does not guarantee a match. The process runs Find, Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Describe the dataset and the de-identification evidence you need at https://sourcex.si/buyers.

Sources

  1. U.S. Department of Health and Human Services, Federal Register (govinfo), "Confidentiality of Substance Use Disorder (SUD) Patient Records (Final Rule), 89 FR, No. 33" (2024). https://www.govinfo.gov/content/pkg/FR-2024-02-16/html/2024-02544.htm
  2. U.S. Department of Health and Human Services, Office for Civil Rights, "Fact Sheet 42 CFR Part 2 Final Rule" (2024). https://www.hhs.gov/hipaa/for-professionals/regulatory-initiatives/fact-sheet-42-cfr-part-2-final-rule/
  3. Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  4. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  5. Accountable, "Addiction Medicine Data Security Requirements: A Practical Guide to HIPAA, 42 CFR Part 2 and EHR Compliance". https://www.accountablehq.com/post/addiction-medicine-data-security-requirements-a-practical-guide-to-hipaa-42-cfr-part-2-and-ehr-compliance
  6. Baker Donelson, "HHS Compliance Deadline Approaching for Updated Part 2 Record Protections". https://www.bakerdonelson.com/hhs-compliance-deadline-approaching-for-updated-part-2-record-protections
  7. Troutman Pepper, "Final Rule Aligns 42 CFR Part 2 With HIPAA/HITECH" (2024). https://www.troutman.com/insights/final-rule-aligns-42-cfr-part-2-with-hipaahitech.html
  8. National Institute of Standards and Technology, "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data