Skip to content

Speech and audio data

Doctor-Patient Conversation Audio for Ambient Clinical Documentation

Quick answer

A usable doctor-patient conversation dataset for ambient scribes is multi-party encounter audio paired with the note the clinician actually signed, plus speaker turns, encounter metadata and documented consent for recording and secondary AI use. Public corpora are small and often simulated, so production teams usually license real encounters. Specify real versus role-played, in-room versus telehealth, channel layout, de-identification method for both audio and text, and how notes were edited after drafting.

By SourceX Editorial · Updated

What separates encounter audio from dictation data

Encounter audio is a conversation between two or more speakers, while dictation is a clinician monologue, and that difference changes nearly every requirement. A scribe model must separate clinician, patient and often a caregiver or interpreter, ignore small talk, and convert lay descriptions ("my chest gets tight on stairs") into structured findings. Dictation data teaches vocabulary and formatting; it does not teach diarization, history taking or note synthesis. For the monologue case, see our guide to physician dictation audio for medical ASR.

The other difference is the target. Dictation pairs audio with a verbatim transcript, whereas an ambient scribe learns a mapping from roughly 10 to 30 minutes of conversation to a note in SOAP or HPI/ROS/Exam/Assessment-and-Plan form. That mapping is lossy and editorial, so the paired note matters as much as the audio.

Public benchmarks are a starting point, not a training set

Public doctor-patient corpora are useful for evaluation baselines but are too small and too synthetic for production training. The research sets that pair conversations with notes typically number in the tens or low hundreds of encounters, and many use simulated visits with actors or clinicians playing patients, because real clinic conversations are rarely recorded and hard to share. Some focus on dialogue snippets rather than full visits.

These sets are valuable for comparing your pipeline against published numbers and for testing note formats. They do not represent real ambient noise, exam-room acoustics, EHR-driven note templates, specialty mix or the way clinicians edit drafts. Treat them as a public sanity check and keep your licensed data as the primary training and held-out evaluation source.

The paired note is the label, so specify which note version you get

The note version you license determines what your model learns, so ask for the final signed note and, where possible, the AI or scribe draft that preceded it. If a supplier already runs an ambient tool, the signed note may be a lightly edited machine draft, and training on it can teach your model another vendor's errors. A draft-plus-final pair is more valuable: the diff shows what clinicians corrected, which is direct supervision for faithfulness.

Fields worth requesting with each note:

  • Note type and template (progress note, H&P, consult, telehealth visit) and specialty.
  • Section boundaries as structured fields, not only a rendered DOCX or PDF.
  • Authorship and edit provenance: human scribe, AI draft, clinician typed, or mixed.
  • Time between encounter end and signature, as a proxy for how much was reconstructed from memory.
  • Codes attached at signature (ICD-10-CM, CPT) if licensed, with the caveat that coding is a separate workflow.

If you also want documentation-gap supervision, clinician query trails are a distinct data type covered in clinical documentation integrity query data.

Audio specification for in-room and telehealth encounters

Recording setup should be specified explicitly because in-room and telehealth audio fail in different ways. In-room capture is usually a single far-field phone or room microphone with overlapping speech, door noise, equipment beeps and a patient who is quieter than the clinician. Telehealth audio is codec-compressed, carries platform echo-cancellation artifacts and, depending on the platform, may be exported as separate per-participant tracks, which makes diarization easier and acoustic realism different.

Ask for:

  • Native sample rate and codec, and whether files were transcoded (WAV/FLAC preferred over MP3 for training; vendors commonly ship MP3 or WAV with JSON or DOCX transcripts [5]). See audio file specs for speech datasets.
  • Channel layout: mono room mix, per-device channels, or separate telehealth tracks.
  • Device and environment labels (smartphone on desk, lapel, ceiling array; exam room, ED bay, home).
  • Speaker roles beyond two: interpreter, family member, nurse, student.
  • Language and code-switching flags, since many visits move between English and another language.

Clinical encounter audio is protected health information until de-identified, and the consent question has two layers: consent to be recorded and permission for secondary AI use. State recording laws vary, and some require every party to agree, so ask how clinician and patient consent were captured, what the consent text said about AI development, and whether participants could withdraw. Role-played encounters with paid actors avoid patient consent but must still be labeled, because models trained on them often underperform on real visits.

HIPAA allows de-identification through Safe Harbor, which removes 18 identifier types, or Expert Determination, in which a qualified expert certifies very small re-identification risk [2]. Neither was written with voice in mind. A voice is itself identifying, so a buyer should ask whether the supplier applied speaker anonymization or treats raw voice as residual risk under an expert determination. The VoicePrivacy Challenge measures exactly this trade-off, scoring anonymized speech for both resistance to speaker verification attacks and remaining ASR utility [3]. Our comparison of HIPAA Safe Harbor and Expert Determination for AI training covers method selection.

Three further traps are specific to encounter data:

  • Substance use disorder discussions may fall under 42 CFR Part 2, which governs confidentiality of SUD treatment records [4]; as of October 2026 its compliance date of 16 February 2026 has passed. Ask whether Part 2 programs were excluded.
  • Spoken identifiers (names, dates of birth, addresses, pharmacy names) must be removed from both audio and transcript, and the two must stay aligned. Bleeps, silence or tags each change what models learn; see how audio redaction affects speech model training.
  • The note text needs its own PHI pass, because notes carry identifiers the audio never mentioned; see de-identifying clinical free text.

Voice data can also raise biometric questions under statutes such as Illinois BIPA; see voiceprints and BIPA. Vendors often state that audio and transcripts were redacted to Safe Harbor guidelines [6]; ask for the method record and a sample check rather than relying on the label. The PHI glossary entry defines the term.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Evaluate the whole pipeline, not only word error rate

An ambient scribe should be evaluated at three stages, because a good transcript can still produce a bad note. The DISPLACE-M challenge for health conversations uses diarization error rate (DER), time-constrained minimum-permutation WER (tcpWER) for speaker-attributed recognition, and ROUGE for downstream summaries [1]. That mirrors a practical stack: who spoke, what they said, and what the note concluded.

ROUGE rewards lexical overlap and misses the failures that matter clinically: a hallucinated medication, a negation flip ("denies chest pain" becoming "chest pain"), or a finding attributed to the wrong speaker. Pair automatic metrics with a clinician rubric scoring omissions, fabrications and attribution errors per note section. For the diarization layer, see evaluating diarization on your own audio and RTTM labels and overlap annotation.

Build a held-out set that is split by clinician and site, not by encounter. If the same clinician appears in train and test, the model can learn that clinician's phrasing and template, inflating scores.

Request template for encounter audio with notes

A precise request shortens assessment because suppliers can check their records against concrete fields. Documentation frameworks such as Data Cards suggest recording sources, collection methods, annotation methods and intended use alongside each dataset [7]; the per-encounter manifest below applies that idea.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "encounter_id": "enc_000418",
  "origin": "real",
  "setting": "in_room",
  "specialty": "family_medicine",
  "visit_type": "follow_up",
  "duration_sec": 1124,
  "audio": {"format": "wav", "sample_rate_hz": 16000, "channels": 1, "device": "smartphone_desk"},
  "speakers": [{"label": "S1", "role": "clinician"}, {"label": "S2", "role": "patient"}, {"label": "S3", "role": "caregiver"}],
  "diarization_file": "enc_000418.rttm",
  "transcript": {"format": "json", "source": "human_verified", "phi_tags": true},
  "note": {"template": "SOAP", "versions": ["ai_draft", "signed_final"], "sections_structured": true},
  "consent": {"recording": "all_parties", "secondary_ai_use": true, "consent_text_version": "v3"},
  "deidentification": {"text_method": "safe_harbor", "audio_method": "spoken_phi_tone_replacement", "voice_anonymized": false, "sample_checked": true},
  "exclusions": ["42_cfr_part_2_programs", "minors"]
}

Decision table for the main choices:

ChoiceOption AOption BWhat to decide before you request
Encounter originReal visitsRole-played with actorsUse real for training and eval; label role-play and keep it out of headline metrics
SettingIn-room far-fieldTelehealth per-participant tracksMatch your deployment; mixing needs a setting label per file
Note versionSigned final onlyAI or scribe draft plus finalPrefer the pair if your model will produce drafts
Voice handlingRaw voice under expert determinationAnonymized voiceAnonymization can reduce ASR utility [3]; test both
TranscriptHuman verifiedASR outputKeep a human-verified subset for evaluation

How SourceX handles requests for clinical encounter data

SourceX sources operational datasets from US companies on request; it does not hold clinical audio in stock, and a request does not guarantee a match. You describe the encounter data you need, not which practices or health systems hold it, and SourceX looks for US businesses that have such records, with every release approved by the supplying company. Its scope also includes new recordings of hands-on work where existing archives do not fit.

Each step follows Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Datasets are rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery. Health records require HIPAA de-identification by Safe Harbor or Expert Determination; personal details are removed or replaced, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement. For broader context, see healthcare administration AI training data, healthcare buyers, the speech and audio data hub, or describe your requirement to SourceX.

Source doctor-patient conversation data for your ambient scribe

If you need encounter audio paired with signed notes, send SourceX the specialty mix, setting, note format, consent and de-identification requirements, and intended uses. SourceX will look for US companies that hold matching records and run the rights review and licensing process with them. Start a buyer request at SourceX.

Sources

  1. arXiv (2603.02813), "Benchmarking Speech Systems for Frontline Health Conversations: The DISPLACE-M Challenge" (2026). https://arxiv.org/pdf/2603.02813
  2. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  3. arXiv (2404.02677), "The VoicePrivacy 2024 Challenge Evaluation Plan" (2024). https://arxiv.org/pdf/2404.02677
  4. U.S. Department of Health and Human Services, "Fact Sheet: 42 CFR Part 2 Final Rule" (2024). https://www.hhs.gov/hipaa/for-professionals/regulatory-initiatives/fact-sheet-42-cfr-part-2-final-rule/
  5. Unidata (vendor page), "Medical Conversations English dataset". https://unidata.pro/datasets/medical-conversations-english/
  6. Shaip (vendor page), "License High-quality Healthcare/Medical Data for AI & ML Models". https://shaip.com/offerings/physician-dictation-audio-data-medical-data-catalog
  7. Pushkarna, Zaldivar, Kjartansson (FAccT 2022; arXiv:2204.01075), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data