Skip to content

Industry-specific operational data

Provider-to-payer phone calls for revenue-cycle voice agents

Quick answer

To train and evaluate a voice agent that calls payers, you need real outbound calls placed by provider or billing-company staff to payer provider-services lines: the full IVR traversal, hold segments, the representative conversation, and the work note or claim outcome that followed. Generic contact-center audio does not cover this. Specify 8 kHz telephony capture, PHI redaction in both audio and transcript, a recording-consent review that covers all-party-consent states, and a join to the resulting accounts-receivable record so every call carries a task-success label.

By SourceX Editorial · Updated

Why provider-to-payer calls are a distinct dataset

Provider-to-payer calls are business-to-business, IVR-heavy conversations between two trained professionals, so they behave differently from inbound consumer support audio. The caller is a biller, AR follow-up specialist or patient-access rep reading a member ID, date of birth, NPI, tax ID and claim number into a payer IVR or to a representative. The patient is usually not on the line, but their identifiers are spoken repeatedly.

Off-the-shelf call-center corpora are typically marketed as general customer-service calls with time-stamped transcripts [2]. They rarely include outbound IVR navigation, DTMF entry, long hold music, transfer loops, or payer-specific vocabulary such as CARC and RARC denial codes, "reprocessing," "corrected claim" or "timely filing." These must be sourced specifically from organizations that place the calls: hospital business offices, physician-group billing teams and RCM outsourcing vendors.

This page focuses on the call itself. For the written AR trail around it, see accounts receivable follow-up histories for claim status agents; for structured 270/271-style results, see eligibility and benefits verification records; and for authorization packets, see prior authorization submissions and payer decisions.

What a usable call record contains

A usable record links audio, transcript, IVR events and the downstream outcome under one call ID. Without the outcome join you can train ASR and turn-taking, but you cannot measure whether the agent actually got the claim status, reference number or authorization it called for.

Ask for these components:

  • Audio: stereo or dual-channel recording (caller and payer on separate channels) at the native telephony rate, plus the recording platform name and codec.
  • IVR event log: prompts heard, DTMF digits or speech responses sent, menu path, timestamps, and where the call reached a live representative or failed out.
  • Segment labels: IVR, hold, transfer, representative, voicemail, and dead air, with start and end offsets.
  • Transcript: time-aligned, speaker-attributed, with PHI spans tagged before redaction so you can audit the masking.
  • Call purpose and payer line: claim status, eligibility and benefits, authorization status, appeal status, or credentialing; payer type (commercial, Medicare Advantage, Medicaid managed care) rather than a named payer if the supplier requires it.
  • Outcome record: the AR or work-queue note written after the call, the reference number given by the representative, and the next action (rebill, appeal, wait, write-off).

The after-call work notes and disposition codes guide covers how to evaluate those notes as summarization targets.

Audio specification: match the 8 kHz channel

Train and evaluate on the channel your agent will actually use, which for payer lines is narrowband telephony. Audio sampled at 8 kHz carries content only up to 4 kHz, and research on telephonic ASR shows gains from pretraining that accounts for the telephone channel rather than assuming wideband speech [1]. Upsampling 8 kHz audio to 16 kHz does not restore the missing band.

Ask the supplier for native sample rate, codec (G.711 mu-law is common on US PSTN legs), whether audio was transcoded by the recording platform, and whether channels were mixed to mono. Mono mixes make diarization of overlapping speech and hold-music removal much harder. Our telephony vs wideband ASR training guide goes deeper, and the contact-center audio requirements template gives a full field list.

These calls carry protected health information even though the patient is absent, so plan redaction and consent review before any audio leaves the supplier. Staff speak member IDs, dates of birth, patient names and claim numbers aloud, and payer reps repeat them back. HIPAA's de-identification standards and its limited data set rules both treat names, phone numbers, medical record numbers and similar direct identifiers as items that must be removed [5]. Redaction has to happen in both the transcript and the audio, typically by aligning tagged transcript spans to timestamps and replacing the audio with silence or tone.

Recording consent has to be checked for both sides of the call, because some states require every party to agree. California prohibits recording a confidential communication without the consent of all parties [6], and payer IVRs often play their own recording notice while provider-side platforms may or may not. Ask how each side was notified. Separately, some states treat voiceprints as biometric identifiers; Texas restricts capturing them for a commercial purpose without notice and consent [7], which matters if you plan speaker embeddings or voice cloning. For the mechanics, see redacting spoken PII from call recordings.

Also confirm who has the right to license the calls. An RCM vendor that placed them on behalf of hospital clients usually needs client authorization; see client data held by service providers.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Labeling for task success, not just transcription

Label each call by whether it achieved its purpose, because a voice agent is judged on the end state, not on word error rate. Agent benchmarks in customer-service domains score whether the final database state matches the goal and measure reliability across repeated trials of the same task, often reported as pass^k [3]. Datasets that pair dialogues with the actions and policy guidelines agents must follow support that kind of scoring [4].

For payer calls, a practical success label is: did the caller obtain the information or commitment needed to take the next AR action, and was it recorded correctly? Failure modes worth labeling separately include IVR dead ends, wrong department transfers, the representative refusing to release information without an extra identifier, disconnected holds, and calls where the stated status contradicts the later 835 remittance.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy it matters
call_idc-000412Joins audio, IVR log, transcript and AR note
call_purposeclaim_statusDefines the success criterion
payer_segmentmedicare_advantageStratifies evaluation without naming the payer
audio8 kHz, G.711, 2 channelsConfirms channel match
ivr_path1 > 3 > [NPI] > [claim #] > 0Trains and tests IVR navigation
hold_seconds1260Tests patience and hold-detection logic
phi_redactiontranscript spans + audio tone, method loggedAudit trail for de-identification
rep_reference_number[REDACTED-REF]Evidence the call produced a usable artifact
outcome_note"Denied CO-16, missing info; corrected claim to be sent"Downstream label for summarization
task_successtruePrimary eval label
consent_basisboth-party notice recordedSupports rights review

Buyer checklist before you sign

Use this list to qualify any supplier of payer call data; a gap in any row is a reason to renegotiate scope rather than proceed.

  • Who placed the calls, and do they hold the rights or need a client's authorization?
  • What recording notice did each party hear, and in which states were callers located?
  • Native sample rate, codec and channel layout, with a short sample before contract.
  • PHI redaction method for audio and text, the residual-risk review, and whether a HIPAA de-identification method was applied.
  • Join keys from call to AR note, claim and remittance, and how many calls lack an outcome.
  • Allowed uses written into the license: ASR training, agent fine-tuning, evaluation, synthetic generation, and whether voice cloning is excluded.

For evaluation-set design, see voice agent evaluation sets with real call scenarios. For adjacent catalogs, see SourceX's call center audio datasets, healthcare revenue cycle datasets, voice agent training data and the explainer can I license call recordings?. The broader industry operational data guide and the AI data hub cover other verticals.

How SourceX approaches provider-to-payer call requests

SourceX sources operational datasets, including support histories and finance workflows, from US companies on request; nothing is held in stock and a request does not guarantee a match. You describe the calls you need, and SourceX looks for US businesses that hold them, with every release approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, and health records require HIPAA de-identification by Safe Harbor or Expert Determination. You can describe your payer-call requirements to SourceX to start that process.

Get payer call recordings for your voice agent

If your team needs provider-to-payer call audio, IVR logs and outcome notes, SourceX can run a request through Find, Assess, Agree, Transact and Manage, with personal details removed or replaced before delivery and terms agreed per deal. Delivery happens through private, access-controlled workflows only after an executed agreement and supplier approval. Start a buyer request.

Sources

  1. arXiv, "Channel-Aware Pretraining of Joint Encoder-Decoder Self-Supervised Model for Telephonic-Speech ASR" (2022). https://arxiv.org/pdf/2211.01669
  2. Unidata, "call center audio". https://unidata.pro/datasets/call-center-audio
  3. arXiv (Sierra Research), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  4. arXiv (Chen et al., NAACL 2021), "Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems" (2021). https://arxiv.org/abs/2104.00783v1
  5. Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  6. California Legislative Information, "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632
  7. Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data