Speech and audio data
Doctor-Patient Conversation Audio for Ambient Clinical Documentation
Quick answer
A usable doctor-patient conversation dataset for ambient scribes is multi-party encounter audio paired with the note the clinician actually signed, plus speaker turns, encounter metadata and documented consent for recording and secondary AI use. Public corpora are small and often simulated, so production teams usually license real encounters. Specify real versus role-played, in-room versus telehealth, channel layout, de-identification method for both audio and text, and how notes were edited after drafting.
By SourceX Editorial · Updated
What separates encounter audio from dictation data
Encounter audio is a conversation between two or more speakers, while dictation is a clinician monologue, and that difference changes nearly every requirement. A scribe model must separate clinician, patient and often a caregiver or interpreter, ignore small talk, and convert lay descriptions ("my chest gets tight on stairs") into structured findings. Dictation data teaches vocabulary and formatting; it does not teach diarization, history taking or note synthesis. For the monologue case, see our guide to physician dictation audio for medical ASR.
The other difference is the target. Dictation pairs audio with a verbatim transcript, whereas an ambient scribe learns a mapping from roughly 10 to 30 minutes of conversation to a note in SOAP or HPI/ROS/Exam/Assessment-and-Plan form. That mapping is lossy and editorial, so the paired note matters as much as the audio.
Public benchmarks are a starting point, not a training set
Public doctor-patient corpora are useful for evaluation baselines but are too small and too synthetic for production training. The research sets that pair conversations with notes typically number in the tens or low hundreds of encounters, and many use simulated visits with actors or clinicians playing patients, because real clinic conversations are rarely recorded and hard to share. Some focus on dialogue snippets rather than full visits.
These sets are valuable for comparing your pipeline against published numbers and for testing note formats. They do not represent real ambient noise, exam-room acoustics, EHR-driven note templates, specialty mix or the way clinicians edit drafts. Treat them as a public sanity check and keep your licensed data as the primary training and held-out evaluation source.
The paired note is the label, so specify which note version you get
The note version you license determines what your model learns, so ask for the final signed note and, where possible, the AI or scribe draft that preceded it. If a supplier already runs an ambient tool, the signed note may be a lightly edited machine draft, and training on it can teach your model another vendor's errors. A draft-plus-final pair is more valuable: the diff shows what clinicians corrected, which is direct supervision for faithfulness.
Fields worth requesting with each note:
- Note type and template (progress note, H&P, consult, telehealth visit) and specialty.
- Section boundaries as structured fields, not only a rendered DOCX or PDF.
- Authorship and edit provenance: human scribe, AI draft, clinician typed, or mixed.
- Time between encounter end and signature, as a proxy for how much was reconstructed from memory.
- Codes attached at signature (ICD-10-CM, CPT) if licensed, with the caveat that coding is a separate workflow.
If you also want documentation-gap supervision, clinician query trails are a distinct data type covered in clinical documentation integrity query data.
Audio specification for in-room and telehealth encounters
Recording setup should be specified explicitly because in-room and telehealth audio fail in different ways. In-room capture is usually a single far-field phone or room microphone with overlapping speech, door noise, equipment beeps and a patient who is quieter than the clinician. Telehealth audio is codec-compressed, carries platform echo-cancellation artifacts and, depending on the platform, may be exported as separate per-participant tracks, which makes diarization easier and acoustic realism different.
Ask for:
- Native sample rate and codec, and whether files were transcoded (WAV/FLAC preferred over MP3 for training; vendors commonly ship MP3 or WAV with JSON or DOCX transcripts [5]). See audio file specs for speech datasets.
- Channel layout: mono room mix, per-device channels, or separate telehealth tracks.
- Device and environment labels (smartphone on desk, lapel, ceiling array; exam room, ED bay, home).
- Speaker roles beyond two: interpreter, family member, nurse, student.
- Language and code-switching flags, since many visits move between English and another language.
Consent, privacy and the rules that attach to this data
Clinical encounter audio is protected health information until de-identified, and the consent question has two layers: consent to be recorded and permission for secondary AI use. State recording laws vary, and some require every party to agree, so ask how clinician and patient consent were captured, what the consent text said about AI development, and whether participants could withdraw. Role-played encounters with paid actors avoid patient consent but must still be labeled, because models trained on them often underperform on real visits.
HIPAA allows de-identification through Safe Harbor, which removes 18 identifier types, or Expert Determination, in which a qualified expert certifies very small re-identification risk [2]. Neither was written with voice in mind. A voice is itself identifying, so a buyer should ask whether the supplier applied speaker anonymization or treats raw voice as residual risk under an expert determination. The VoicePrivacy Challenge measures exactly this trade-off, scoring anonymized speech for both resistance to speaker verification attacks and remaining ASR utility [3]. Our comparison of HIPAA Safe Harbor and Expert Determination for AI training covers method selection.
Three further traps are specific to encounter data:
- Substance use disorder discussions may fall under 42 CFR Part 2, which governs confidentiality of SUD treatment records [4]; as of October 2026 its compliance date of 16 February 2026 has passed. Ask whether Part 2 programs were excluded.
- Spoken identifiers (names, dates of birth, addresses, pharmacy names) must be removed from both audio and transcript, and the two must stay aligned. Bleeps, silence or tags each change what models learn; see how audio redaction affects speech model training.
- The note text needs its own PHI pass, because notes carry identifiers the audio never mentioned; see de-identifying clinical free text.
Voice data can also raise biometric questions under statutes such as Illinois BIPA; see voiceprints and BIPA. Vendors often state that audio and transcripts were redacted to Safe Harbor guidelines [6]; ask for the method record and a sample check rather than relying on the label. The PHI glossary entry defines the term.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Evaluate the whole pipeline, not only word error rate
An ambient scribe should be evaluated at three stages, because a good transcript can still produce a bad note. The DISPLACE-M challenge for health conversations uses diarization error rate (DER), time-constrained minimum-permutation WER (tcpWER) for speaker-attributed recognition, and ROUGE for downstream summaries [1]. That mirrors a practical stack: who spoke, what they said, and what the note concluded.
ROUGE rewards lexical overlap and misses the failures that matter clinically: a hallucinated medication, a negation flip ("denies chest pain" becoming "chest pain"), or a finding attributed to the wrong speaker. Pair automatic metrics with a clinician rubric scoring omissions, fabrications and attribution errors per note section. For the diarization layer, see evaluating diarization on your own audio and RTTM labels and overlap annotation.
Build a held-out set that is split by clinician and site, not by encounter. If the same clinician appears in train and test, the model can learn that clinician's phrasing and template, inflating scores.
Request template for encounter audio with notes
A precise request shortens assessment because suppliers can check their records against concrete fields. Documentation frameworks such as Data Cards suggest recording sources, collection methods, annotation methods and intended use alongside each dataset [7]; the per-encounter manifest below applies that idea.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"encounter_id": "enc_000418",
"origin": "real",
"setting": "in_room",
"specialty": "family_medicine",
"visit_type": "follow_up",
"duration_sec": 1124,
"audio": {"format": "wav", "sample_rate_hz": 16000, "channels": 1, "device": "smartphone_desk"},
"speakers": [{"label": "S1", "role": "clinician"}, {"label": "S2", "role": "patient"}, {"label": "S3", "role": "caregiver"}],
"diarization_file": "enc_000418.rttm",
"transcript": {"format": "json", "source": "human_verified", "phi_tags": true},
"note": {"template": "SOAP", "versions": ["ai_draft", "signed_final"], "sections_structured": true},
"consent": {"recording": "all_parties", "secondary_ai_use": true, "consent_text_version": "v3"},
"deidentification": {"text_method": "safe_harbor", "audio_method": "spoken_phi_tone_replacement", "voice_anonymized": false, "sample_checked": true},
"exclusions": ["42_cfr_part_2_programs", "minors"]
}
Decision table for the main choices:
| Choice | Option A | Option B | What to decide before you request |
|---|---|---|---|
| Encounter origin | Real visits | Role-played with actors | Use real for training and eval; label role-play and keep it out of headline metrics |
| Setting | In-room far-field | Telehealth per-participant tracks | Match your deployment; mixing needs a setting label per file |
| Note version | Signed final only | AI or scribe draft plus final | Prefer the pair if your model will produce drafts |
| Voice handling | Raw voice under expert determination | Anonymized voice | Anonymization can reduce ASR utility [3]; test both |
| Transcript | Human verified | ASR output | Keep a human-verified subset for evaluation |
How SourceX handles requests for clinical encounter data
SourceX sources operational datasets from US companies on request; it does not hold clinical audio in stock, and a request does not guarantee a match. You describe the encounter data you need, not which practices or health systems hold it, and SourceX looks for US businesses that have such records, with every release approved by the supplying company. Its scope also includes new recordings of hands-on work where existing archives do not fit.
Each step follows Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Datasets are rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery. Health records require HIPAA de-identification by Safe Harbor or Expert Determination; personal details are removed or replaced, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement. For broader context, see healthcare administration AI training data, healthcare buyers, the speech and audio data hub, or describe your requirement to SourceX.
Source doctor-patient conversation data for your ambient scribe
If you need encounter audio paired with signed notes, send SourceX the specialty mix, setting, note format, consent and de-identification requirements, and intended uses. SourceX will look for US companies that hold matching records and run the rights review and licensing process with them. Start a buyer request at SourceX.
Sources
- arXiv (2603.02813), "Benchmarking Speech Systems for Frontline Health Conversations: The DISPLACE-M Challenge" (2026). https://arxiv.org/pdf/2603.02813
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- arXiv (2404.02677), "The VoicePrivacy 2024 Challenge Evaluation Plan" (2024). https://arxiv.org/pdf/2404.02677
- U.S. Department of Health and Human Services, "Fact Sheet: 42 CFR Part 2 Final Rule" (2024). https://www.hhs.gov/hipaa/for-professionals/regulatory-initiatives/fact-sheet-42-cfr-part-2-final-rule/
- Unidata (vendor page), "Medical Conversations English dataset". https://unidata.pro/datasets/medical-conversations-english/
- Shaip (vendor page), "License High-quality Healthcare/Medical Data for AI & ML Models". https://shaip.com/offerings/physician-dictation-audio-data-medical-data-catalog
- Pushkarna, Zaldivar, Kjartansson (FAccT 2022; arXiv:2204.01075), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.