Fine-tuning and post-training data
Healthcare LLM fine-tuning data: clinical vs administrative records under HIPAA
Quick answer
A medical LLM fine-tuning dataset usually comes from one of two record families. Clinical records (progress notes, discharge summaries, radiology reports) teach medical reasoning but carry the most identifiers and the hardest de-identification work. Administrative records (prior authorization requests, denial letters, appeal outcomes, claim edits) carry lower clinical risk and often contain verifiable targets. Either way, licensed US records should arrive de-identified under HIPAA Safe Harbor or Expert Determination, with the method documented [1][2].
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why public medical instruction sets are not enough
Public medical instruction sets mostly avoid real clinical records, so they teach exam-style knowledge rather than the documentation patterns your product will see. MedAlpaca's Medical Meadow, for example, assembles 160,000+ entries from sources such as WikiDoc question-answer pairs and licensing-exam style tasks [6]. ChatDoctor fine-tunes on about 100k patient-physician conversations drawn from online consultation material, and its authors argue for on-premises models to protect patient data [7].
Those sets are useful for warm-up and knowledge coverage. They do not contain a hospitalist's abbreviation-heavy progress note, a payer's denial rationale citing a medical policy number, or an appeal letter that overturned it. If your model drafts prior authorization letters or summarizes inpatient stays, you need examples drawn from the operational records that generate those artifacts. Check the license on any public set before commercial use; a large audit of dataset hosting sites found license omission above 70% and license error rates above 50% [10]. See open instruction and preference datasets that allow commercial fine-tuning for a closer look.
Clinical versus administrative records: value, risk and process
Clinical text gives the richest supervision for medical reasoning, while administrative records give cleaner, more checkable targets at lower identification risk. The table below compares the two families for fine-tuning work.
| Dimension | Clinical records (notes, reports, summaries) | Administrative records (prior auth, denials, appeals, claims) |
|---|---|---|
| Typical SFT tasks | Note summarization, problem-list extraction, patient-message drafting | Letter drafting, denial-reason classification, criteria matching, coding-to-policy extraction |
| Verifiable targets | Weak; "correct" summary is a judgment call | Stronger; outcome fields (approved, denied, overturned) can grade outputs |
| Preference and RL signal | Requires clinician raters, which are scarce and costly | Appeal outcomes and reviewer edits act as natural preference pairs |
| Identifier density | High in free text: names, dates, facility names, family members | Moderate: member IDs, NPIs, dates of service, addresses on letters |
| De-identification route | Expert Determination usually needed to keep dates and context useful | Safe Harbor often workable when year-only dates are acceptable |
| Template overfitting risk | EHR note templates vary by vendor and site | Payer letter templates are highly repetitive within one organization |
Free-text notes are harder to de-identify than structured fields because identifiers appear in unpredictable places: "pt's daughter Maria called from Tucson" defeats a field-level scrubber. NIST's survey of de-identification documents how supposedly de-identified data has been re-identified, which is why a recorded method and a sample check matter more than a vendor's assurance [5]. For administrative AI use cases that do not involve fine-tuning design, the owner page is healthcare administration AI training data. For payer review files specifically, see utilization management review records.
Which HIPAA de-identification route a fine-tuning buyer should require
Require Safe Harbor when your tasks survive losing all date elements finer than year and most geography; require Expert Determination when temporal sequence or regional context drives the task. Under Safe Harbor, 18 categories of identifiers of the individual and of relatives, employers and household members are removed, and the covered entity must have no actual knowledge that the remainder could identify someone [1][2]. Expert Determination instead relies on a qualified expert applying statistical and scientific principles to conclude that re-identification risk is very small, documented in a report [1].
For fine-tuning, the practical difference is dates. A discharge-summary model trained on Safe Harbor data sees years only, so "post-op day 3" logic must come from relative phrasing in the text. Expert Determination can allow consistent date shifting per patient, which preserves intervals. Detailed comparisons are on HIPAA Safe Harbor vs Expert Determination for AI training and dates, ages and ZIP codes under Safe Harbor. If you receive an expert report, use the checklist in how to review a HIPAA Expert Determination report.
A limited data set is not de-identified. It removes 16 direct identifiers but may keep dates and some geography, and it can only be used under a data use agreement for research, public health or health care operations [2]. Do not accept a limited data set described as "HIPAA compliant" training data without counsel review. Substance use disorder treatment records add 42 CFR Part 2; the 2024 final rule aligned Part 2 more closely with HIPAA, including on de-identification, and as of October 2026 its compliance date of 16 February 2026 has passed [11].
Who may touch the data before de-identification
Business associate agreements and data use agreements decide who may process identifiable records, so confirm the chain before any vendor touches raw notes. HHS states that a business associate may use or disclose PHI only as its contract permits or as required by law [4]. A business associate may de-identify PHI only if its agreement authorizes that [3].
In practice, ask three questions. Who ran de-identification, and under what agreement? Did any annotation or labeling vendor see records before de-identification, and if so, under a BAA? Is the de-identification method and the date it was applied written into the diligence file? A dataset de-identified by a party with no authority to do so is a provenance problem even if the output is clean. The HIPAA-compliant AI training data page explains what that label must mean.
Designing SFT and preference examples from healthcare records
Good healthcare fine-tuning examples pair a real input document with an output a clinician or reviewer actually produced or approved. Small, carefully curated sets can carry much of the alignment effect: LIMA fine-tuned a 65B model on 1,000 curated prompt-response pairs and performed competitively in human comparisons [8]. That favors a few thousand high-quality, reviewer-checked healthcare pairs over a large pool of scraped Q&A.
Useful pair types include: prior authorization request to medical-necessity determination; denial letter to structured fields (reason code, policy reference, missing documentation); appeal packet to outcome with reviewer rationale; and discharge summary to patient-friendly instructions. Revision history is valuable for preference tuning, since an approved final letter versus the first draft gives a chosen/rejected pair without new labeling. Format guidance is in structured-output fine-tuning data and chat fine-tuning data format.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"record_family": "administrative",
"task": "denial_reason_extraction",
"deid_method": "safe_harbor",
"deid_applied": "2026-08-14",
"source_site_id": "site_07",
"input": "Request for lumbar MRI denied. Documentation does not show 6 weeks of conservative therapy per policy [POLICY_ID]. Member [MEMBER_ID] may appeal within [N] days.",
"output": {
"decision": "denied",
"reason_category": "insufficient_conservative_therapy",
"policy_reference": "[POLICY_ID]",
"missing_documentation": ["physical therapy notes", "duration of symptoms"]
},
"preference": {"chosen": "final_reviewer_version", "rejected": "first_draft"},
"reviewer_role": "utilization_review_nurse"
}
Evaluating across sites and payers
Hold out whole organizations, not random rows, or your evaluation will reward memorized templates. Healthcare documents are heavily templated: one health system's discharge summaries share section headers and smart-phrases, and one payer's denial letters reuse paragraphs. A random split leaks those templates into the test set and inflates scores.
Split by source_site_id or payer, and keep at least one organization entirely unseen. Run a memorization check as well, because training text can be extracted from language models by querying them, and residual identifiers in free text are exactly what you do not want reproduced [9]. Pre-purchase checks are covered in how to evaluate a fine-tuning dataset before you buy it, and privacy-preserving training options in differential privacy for LLM fine-tuning.
Request checklist for healthcare fine-tuning data
Describe the data and the task precisely; suppliers can only assess a request they understand.
- Record family and document types (for example, denial letters plus appeal outcomes, or inpatient discharge summaries).
- Target tasks: SFT, preference pairs, structured extraction, or a mix.
- Required de-identification route and which fields you need preserved (relative dates, age bands, state).
- Whether revision history or reviewer edits must be included.
- Number of distinct organizations needed for a site-held-out evaluation.
- Exclusions: Part 2 records, psychotherapy notes, pediatric records, or anything your counsel rules out.
- Delivery format (JSONL, FHIR resources, PDF plus extracted text) and annotation needs.
How SourceX helps healthcare AI teams
SourceX sources operational datasets from US companies on request, including documents and finance and legal workflows, and manages licensing and ongoing purchases; nothing is held in stock, and a request does not guarantee a match. Health records require HIPAA de-identification by Safe Harbor or Expert Determination, every dataset is rights-reviewed and licensed with defined records, uses, term and delivery, and each release is approved by the supplying company. You can describe the healthcare records you need on the SourceX buyers page; related context sits on healthcare buyers, licensing medical records for AI training and HIPAA and AI training data.
Source de-identified healthcare records for fine-tuning
SourceX looks for US businesses that hold the healthcare records you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is contracted. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Start a healthcare data request on the SourceX buyers page.
Frequently asked questions
Can I fine-tune on a limited data set instead of de-identified data?
A limited data set is still PHI and may be used only under a data use agreement for research, public health or health care operations [2]. Most commercial model training does not fit those purposes cleanly, so get counsel review before relying on it.
Are administrative records lower risk than clinical notes?
Generally yes for clinical sensitivity, but letters still contain member IDs, addresses and dates of service. They still need full de-identification, and their repetitive templates raise overfitting risk.
How many examples do I need?
There is no fixed number. Curated quality matters more than volume, as the LIMA results suggest [8]; size the set by task coverage and the number of held-out organizations you need. More on scoping is in how to source supervised fine-tuning data and the fine-tuning buyer's guide.
Sources
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- U.S. Department of Health and Human Services, "May a health information organization (HIO), acting as a business associate of a HIPAA covered entity, de-identify information?". https://www.hhs.gov/hipaa/for-professionals/faq/544/may-a-health-information-organization-de-identify-information/index.html
- U.S. Department of Health and Human Services, "Business Associate Contracts (sample business associate agreement provisions)". https://www.hhs.gov/hipaa/for-professionals/covered-entities/sample-business-associate-agreement-provisions/index.html
- National Institute of Standards and Technology, "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf
- arXiv (Han et al.), "MedAlpaca: An Open-Source Collection of Medical Conversational AI Models and Training Data" (2023). https://arxiv.org/pdf/2304.08247
- arXiv (Li et al.), "ChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain Knowledge" (2023). https://arxiv.org/pdf/2303.14070v4
- arXiv (Zhou et al.), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- arXiv (Nasr et al.), "Scalable Extraction of Training Data from (Production) Language Models" (2023). https://arxiv.org/abs/2311.17035v1
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Troutman Pepper, "Final Rule Aligns 42 CFR Part 2 with HIPAA/HITECH" (2024). https://www.troutman.com/insights/final-rule-aligns-42-cfr-part-2-with-hipaahitech.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.