Skip to content

Industry-specific operational data

Patient portal message data for in-basket triage and reply drafting AI

Quick answer

A usable patient portal messages dataset pairs each inbound patient message with what the practice actually did: the category and urgency it was routed under, the pool or role that handled it, the reply sent, and timing. For triage classifiers you need the routing and escalation labels; for reply-drafting SFT and evaluation you need message-reply pairs with a flag showing which replies were AI-assisted. All of it is protected health information until de-identified under HIPAA, so the de-identification method and license terms matter as much as volume.

By SourceX Editorial · Updated

What a portal message record should contain

A training-grade record is a thread, not a single message: the patient's original text, every staff and clinician turn, the routing history and the closing action. Health systems on Epic MyChart, Oracle Health (Cerner) HealtheLife or athenahealth's patient portal store these as in-basket or task objects with pool assignments, so ask for an export that preserves thread order and handoffs rather than flattened text.

The fields that carry the signal are:

  • Thread and turn IDs with timestamps, so you can rebuild the conversation and compute time to first reply and time to close.
  • Sender role on each turn: patient, proxy (parent or caregiver), front desk, medical assistant, RN, pharmacist, APP or physician.
  • Routing history: the initial pool, every reassignment and the final owner. Reassignments are free labels for "the first triage was wrong."
  • Patient-selected topic from the portal form (for example "Medication question" or "Test results") next to the staff-applied category, because the two often disagree.
  • Closing action: reply only, refill sent, appointment booked, e-visit billed, phone call placed, or redirected to urgent care or the emergency department.
  • Linked orders or encounters, coded rather than free text, when the message triggered a refill, a lab order or a visit.

The medical records guide covers notes and charts; portal messages are a different asset because they are asynchronous, written by patients in lay language and closed by an operational action.

Labels that make triage and urgency models work

Triage models learn from the decisions staff actually made, so the most useful labels come from the workflow itself rather than from fresh annotation. A practical taxonomy separates category (refill, results question, scheduling, clinical question, billing, forms and letters, administrative) from urgency (routine, same-day, escalated to phone, redirected to emergency care) and from responder role.

Urgency is the hardest label to source. Most portals display a "do not use for emergencies" banner, so true emergencies are rare and handled by phone, which makes the escalated class tiny and noisy. Ask suppliers how escalation was recorded: a structured flag, a routing to a nurse triage pool, or only a phrase in the reply such as "please call 911." If it is only in text, budget for a labeling pass and keep an adjudicated gold subset for evaluation.

Two operational labels are worth requesting even if you do not train on them. Time to reply supports workload and SLA modeling, and billing metadata from e-visits (for example the online digital evaluation CPT codes 99421-99423 recorded when a message is converted to a billable service) shows which messages needed clinical judgment. For general methods on routing labels, see classification fine-tuning data.

Separating human replies from AI-drafted replies

Since health systems began piloting LLM reply drafting in the in-basket, a share of recent replies may have started as model drafts, so any reply corpus from the drafting era needs a provenance flag. Several published health-system pilots have measured draft use rates and blinded clinician ratings of AI versus human replies; read them directly for figures before assuming how much of a supplier's corpus is affected.

Without a flag you risk training a drafting model on its predecessor's output and inheriting its tone and omissions. Ask whether the source system logged a draft-used or draft-edited indicator, and whether the original draft text was retained so you can compute the clinician's edit. Draft-to-final edit pairs are themselves a strong preference and evaluation signal.

For evaluation sets, keep every reference reply human-authored or human-approved and record which is which. That keeps your gold standard independent of any deployed drafting tool.

De-identifying free-text patient messages

Patient messages leak identifiers in places structured de-identification never looks, so a free-text NLP pipeline plus sampled human review is the baseline. HIPAA allows two routes: Safe Harbor, which removes 18 identifier types including those of relatives, employers and household members, with no actual knowledge that the remainder could identify the patient, or Expert Determination by a person with appropriate statistical and scientific expertise [1][2].

The failure modes are specific to this channel. Patients sign with family names, mention a spouse's employer, paste pharmacy addresses, include dates of birth in refill requests and attach photos of rashes, insurance cards or pill bottles. Proxy messages from parents name the child and the parent. Free-text de-identification needs NER-based PHI detection tuned on clinical text and a documented sample check [6].

No method catches everything: manual and automated removal both leave residual identifiers, and replacing detected PHI with realistic surrogates ("hiding in plain sight") makes the leftovers hard to distinguish [3]. Re-identification of supposedly de-identified data is well documented [4], so treat the delivered set as sensitive regardless of its label. See HIPAA-compliant AI training data for what that label must mean before you license.

Two further cases need explicit handling:

  • Attachments. Images and PDFs usually cannot be cleared by text pipelines. Most buyers exclude them or receive only a type tag (photo, document) and a caption-free placeholder.
  • Substance use disorder content. Messages from a federally assisted SUD program fall under 42 CFR Part 2; the 2024 final rule aligned Part 2's de-identification standard with HIPAA [5], with a compliance date of 16 February 2026, so confirm the supplier screened for Part 2 records.

If you need dates or ZIP-level geography for timing analysis, a limited data set under a data use agreement (164.514(e)) is an alternative to full de-identification, with different contractual obligations [2].

Request template for a patient message dataset

The fastest way to get a meaningful answer from a supplier is a request that names the fields, labels, time window and de-identification method up front.

Illustrative example: invented to show structure; it does not describe an available dataset.

ItemWhat to specify
Use caseTriage classifier, reply-drafting SFT, evaluation set, or all three
Specialty mixPrimary care, oncology, cardiology, pediatrics (proxy-heavy), behavioral health (Part 2 risk)
Time windowDates before and after any AI drafting rollout, with the rollout date
Thread scopeFull threads with all staff turns, routing history and closing action
LabelsCategory, urgency or escalation, responder role, time to reply, e-visit billing flag
ProvenanceAI-draft used or edited flag; original draft text if retained
ExclusionsAttachments, Part 2 records, messages from minors' own accounts if required
De-identificationSafe Harbor or Expert Determination, method report, sample-check results
FormatJSONL or Parquet, one row per turn, with a Croissant metadata file [7]
LicenseAllowed uses (training, evaluation), term, delivery and deletion terms

An illustrative record for one turn:

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "thread_id": "T-000184",
  "turn": 2,
  "sender_role": "RN",
  "patient_topic": "Medication question",
  "staff_category": "clinical_question",
  "urgency": "same_day",
  "routed_from_pool": "front_desk",
  "routed_to_pool": "rn_triage",
  "text": "Thanks for letting us know about the swelling in [BODY_PART]. Please call the clinic today so we can ...",
  "ai_draft_used": true,
  "ai_draft_edited": true,
  "minutes_since_prior_turn": 142,
  "closing_action": "phone_call_placed",
  "deid_method": "expert_determination"
}

Evaluating a clinician inbox model on real messages

A credible clinician inbox evaluation holds out recent, human-replied threads and scores the model against what the care team actually did. For triage, report per-class recall with special weight on the escalated class, since a missed urgent message is the costly error. For drafting, compare against human references on clinical accuracy, completeness, safety (did it advise appropriate escalation) and tone, using blinded clinician raters who do not know which reply came from the model.

Calibrate any automated grader against those clinician ratings before trusting it at scale; see LLM-as-a-judge calibration sets. Outcome labels such as "refill sent" or "visit booked" provide ground truth that does not depend on raters, as covered in outcome-labeled evaluation data.

Split by time, not randomly, so threads from the same patient and period do not leak across train and test. Keep a separate slice from after the AI-drafting rollout to measure whether your model simply imitates the deployed one.

How licensing and sourcing typically work

Portal message data comes from provider organizations, and access is negotiated per deal rather than bought off a shelf. Buyers usually need the covered entity's approval, a documented de-identification method, and a license that states which records, uses, term and delivery apply. Related patient-access data, such as referral and scheduling records, often comes from the same operational teams and can be scoped in the same conversation.

SourceX sources operational datasets from US companies on request; categories are not inventory and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, and health records require HIPAA de-identification by Safe Harbor or Expert Determination. Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Buyers in healthcare can see how this applies on the healthcare buyers page or start a request through SourceX for AI data buyers. Teams sourcing general support conversations rather than clinical ones may also find the chat logs page relevant, and the industry-specific operational data hub lists neighboring categories.

Source patient message data for your triage or drafting model

SourceX looks for US healthcare organizations that hold the message data you describe, reviews data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Every release is approved by the supplying organization, and delivery runs through private, access-controlled workflows after an executed agreement. Describe the threads, labels and de-identification you need at SourceX for AI data buyers.

Frequently asked questions

Can de-identified portal messages still be traced to a patient?

Yes, in some cases. Residual identifiers survive both manual and automated scrubbing [3], and de-identified data has been re-identified in documented cases [4]. Treat delivered messages as sensitive, restrict access, and prohibit re-identification in the license.

Are proxy messages from parents or caregivers usable?

They are often the richest pediatric and geriatric data, but they contain identifiers for two people. Confirm the de-identification pipeline handles both the proxy and the patient, and check any rules on adolescent confidential messaging before including them.

Should we exclude message attachments?

Usually, unless you are building an image model and can apply a separate review. Photos of skin, documents and insurance cards carry faces, names and member numbers that text pipelines miss. This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Sources

  1. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  2. Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  3. Journal of the American Medical Informatics Association (Carrell et al.), "Hiding in plain sight: use of realistic surrogates to reduce exposure of protected health information in clinical text" (2013). https://academic.oup.com/jamia/article/20/2/342/897812
  4. National Institute of Standards and Technology, "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf
  5. Troutman Pepper, "Final Rule Aligns 42 CFR Part 2 With HIPAA/HITECH" (2024). https://www.troutman.com/insights/final-rule-aligns-42-cfr-part-2-with-hipaahitech.html
  6. Ertas, "HIPAA-compliant AI training data guide". https://www.ertas.ai/blog/hipaa-compliant-ai-training-data-guide
  7. MLCommons Croissant working group, "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data