Industry-specific operational data
Referral management and scheduling records for patient access AI
Quick answer
Patient scheduling data for AI is useful only when four layers arrive together: referral intake (often fax images plus the fields staff keyed from them), scheduling event logs, the schedule templates and booking rules that constrained each decision, and outcomes such as arrival, no-show, late cancellation and referral closure. Logs without rules teach an agent the wrong constraints. Expect HIPAA de-identification to decide whether appointment timing survives, so settle that method before scoping volume.
By SourceX Editorial · Updated
What a usable patient access dataset contains
A usable dataset links each referral to every scheduling event it produced and to the rule set in force at that moment. Most providers hold these in different places: a referral work queue in the EHR or a referral management tool, fax server archives, the scheduling module's audit trail, and template build tables maintained by a schedule-build team. Buyers who request only the appointment table get the end state, not the decisions.
The four layers, and what each one trains:
- Referral orders and intake. Inbound referral document (fax TIFF or PDF, eReferral, or HL7/FHIR message), received timestamp, referring provider and specialty, ordered service, diagnosis codes, urgency, required attachments, missing-information flags, and the work-queue status history. Trains document extraction and intake triage.
- Scheduling event logs. Every book, reschedule, cancel, bump and waitlist action, with actor role (patient portal, call center, clinic staff, bot), reason code, offered slots versus chosen slot, and lead time. Trains scheduling agents and next-best-action models.
- Schedule templates and rules. Provider templates, visit types and durations, slot reservations (new-patient holds, urgent holds), double-booking limits, location and equipment constraints, and provider preferences written as free text. Trains constraint handling.
- Outcomes. Arrived, no-show, late cancellation, same-day reschedule, time from referral to first appointment, referral closed-loop status, and whether care left the network (leakage). Provides labels.
Neighboring workflows have their own pages: eligibility checks that usually precede booking are covered in insurance eligibility and benefits verification records, and payer medical-necessity review sits in utilization management review records.
Why schedule templates matter more than booking logs
Schedule templates matter because the rules, not the bookings, explain why a slot was or was not offered. A log showing that a new-patient cardiology visit was booked 23 days out looks like a preference unless the template shows that new-patient slots were held only on Tuesday afternoons. An agent trained on logs alone learns the artifacts of a template it never saw.
The HL7 FHIR data model reflects the same separation. In FHIR R4, a Schedule is a container that defines the period and the appointment types that can be booked, and it carries no appointment information [5]. Slots are bookable intervals that can hold more than one appointment, can be marked busy with no appointment attached (blocked time) [5], and carry no recurrence information [5], which is managed outside the resource. Recurrence and blocking logic therefore often lives in vendor-specific template tables, not in exported FHIR data.
Ask for template versions with effective dates, so each booking can be joined to the rules active on that date. Also ask for free-text provider preference notes ("no new patients after 3pm", "MRI must precede consult"); they are messy but they are the actual constraints. Template text repeated across thousands of records also needs handling, as described in templates and boilerplate in business records.
Fax referral intake for document extraction
Fax referral data is valuable when each document image is paired with the fields staff actually entered and the corrections they made. The image alone has no labels; the work-queue record alone has no input. The pair is a supervised extraction example.
Request page-level images, the indexed document type (referral letter, clinical notes, insurance card, imaging order), the keyed fields, and the status history showing "incomplete, sent back for records" events. Those send-back events are the negative examples an intake model needs. Check how often staff overwrote extracted values, because high override rates indicate noisy labels; see verifying outcome labels in operational records.
No-show, leakage and closure labels
Outcome labels are reliable only when you know how the source system assigns them. A no-show flag may be set automatically at end of day, manually by front-desk staff, or not at all for telehealth visits; late cancellations may be coded as cancellations, no-shows or a separate status depending on clinic policy. Ask for the status dictionary and the rule that sets each value.
Leakage and closed-loop labels are harder. Closure usually requires a consult note or result returned to the referring provider; leakage is often inferred from claims or simply recorded as "patient chose another provider". Ask whether the label is observed or inferred, and what fraction of referrals has no terminal status.
HIPAA de-identification and appointment timing
The de-identification method determines whether timing patterns survive. Under the HIPAA Safe Harbor method, all elements of dates (except year) directly related to an individual must be removed, along with the other listed identifiers [1][2]. Appointment dates, referral dates and admission dates fall in that category, which strips the day-of-week, lead-time and seasonality signals that no-show and capacity models depend on.
Expert Determination lets a qualified expert certify that re-identification risk is very small for a specific dataset and recipient, which can keep shifted dates or derived intervals at the cost of more analysis and documentation [1][3]. A limited data set under a data use agreement is a third route that can retain dates [2]. Practical options include per-patient date shifting, converting dates to intervals (days from referral to booking, booking to visit), and generalizing ZIP codes. Referrals into substance use disorder programs may also fall under 42 CFR Part 2, whose amended rule had a compliance date of February 16, 2026 [4]; ask suppliers to flag or exclude those programs.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Fairness fields for no-show models
No-show models need fairness review because predicted risk can track socioeconomic status and race, and the harm often appears in how predictions are used. Published operations research has reported that when a group carries higher predicted no-show risk, overbooking policies can place those patients into overbooked slots more often, lengthening their waits [6]. Features like distance, payer class, prior no-shows and portal use can act as proxies.
Ask for de-identified demographic or area-level fields usable for subgroup evaluation, not as model inputs, plus overbooking decisions and wait-time outcomes so you can measure downstream allocation. A structured approach is in the dataset bias audit guide.
Request template for referral and scheduling data
A clear request names the layers, join keys, de-identification method and labels up front.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field group | Example fields | Format | Why it matters |
|---|---|---|---|
| Referral | referral_id, received_ts (shifted), source_channel (fax/eReferral/portal), specialty, service_code, urgency, missing_info_flags | Parquet + page images (TIFF/PDF) | Intake extraction and triage |
| Work queue | referral_id, status, status_ts, actor_role, sendback_reason | Parquet | Negative examples, cycle time |
| Schedule events | appt_id, referral_id, event_type (book/reschedule/cancel/bump), offered_slot_ids, chosen_slot_id, reason_code, actor_role | Parquet | Agent decision traces |
| Templates | template_id, provider_key, visit_type, duration_min, hold_type, max_overbook, effective_from/to, preference_text | JSON | Constraint learning |
| Outcomes | appt_id, final_status, status_rule, days_referral_to_visit, closed_loop, leakage_flag (observed/inferred) | Parquet | Labels |
| Governance | deid_method, date_shift_policy, excluded_programs, subgroup_fields | Data card | Review and fairness checks |
For decision traces generally, see decision records with rationale, and for proof that a provider can license the records, see chain of title for AI training data.
How SourceX handles patient access data requests
SourceX sources operational datasets from US companies on request; it does not hold stock, and a request does not guarantee a match. Buyers describe the data, not the businesses, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, health records require HIPAA de-identification by Safe Harbor or Expert Determination, and delivery runs through private, access-controlled workflows only after an executed agreement. Start from the buyer request page, or see the healthcare administration buyer overview, healthcare buyers and healthcare administration AI training data. Other operational categories are listed in the industry data hub and the AI data hub.
Request patient scheduling data for AI
SourceX looks for US providers that hold referral, scheduling and outcome records like those described here, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything transacts. Nothing is contracted until a supplier agrees. Describe the scheduling data you need.
Sources
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- Censinet, "Safe Harbor vs Expert Determination for PHI". https://censinet.com/perspectives/safe-harbor-vs-expert-determination-phi
- U.S. Department of Health and Human Services (SAMHSA and OCR), Federal Register, "Confidentiality of Substance Use Disorder (SUD) Patient Records, Final Rule" (2024). https://www.govinfo.gov/content/pkg/FR-2024-02-16/html/2024-02544.htm
- HL7 International, "Slot - FHIR v4.0.1". https://www.hl7.org/fhir/R4/slot.html
- Samorani, M., et al., "Overbooked and Overlooked: Machine Learning and Racial Bias in Medical Appointment Scheduling" (2021). https://pubsonline.informs.org/doi/10.1287/msom.2021.0999
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.