Industry-specific operational data
Accounts receivable follow-up histories for claim status and collector agents
Quick answer
An AR follow-up dataset for a claim status agent is the full work history of unpaid claims: each 276/277 status check, payer portal or phone touch, collector note, action code, next follow-up date and final resolution with dollars collected or written off. Buy complete sequences to resolution, with the supplier's action-code glossary and payer-level context, not isolated notes. Without the sequence and the glossary, an agent can learn note style but not which next action actually gets a claim paid.
By SourceX Editorial · Updated
This page covers the multi-touch trajectory: what a collector did on day 31, 45 and 70 of a claim and why. Single-label denial prediction from matched claims and remittances is a different problem, covered in claim denial prediction training data, and recorded payer calls are covered in provider-to-payer phone calls for revenue-cycle voice agents. For the wider category, see the industry-specific operational data hub and SourceX's healthcare revenue cycle datasets.
What an AR follow-up trajectory contains
A usable trajectory is an ordered event log keyed to one claim, from the first aged-AR trigger to a terminal state. In most practice-management and RCM platforms, collectors work from a queue filtered by aging bucket (31-60, 61-90, 91-120, 120+ days), payer and balance, and every touch writes an activity row. The training value sits in the ordering: status returned, action taken, note written, follow-up date set, and what happened next.
The core event types to require are:
- Status inquiries and responses. Electronic 276 requests and 277 responses in the ASC X12 005010X212 format, which HIPAA names for claim status, with CAQH CORE operating rules adopted for that transaction since January 1, 2013. Keep the 277 STC status category and status codes, not just a "pending" flag.
- Portal and phone touches. Payer portal lookups and calls, with the call reference number, the representative's stated reason and any promised reprocessing date.
- Collector actions. Coded actions such as rebill, corrected claim, submit records, appeal, transfer to patient responsibility, adjust or escalate, each with user role and timestamp.
- Remittance and adjustment events. 835 payments and adjustments with group codes and claim adjustment reason codes, whose official list X12 maintains and revises, with payer guidance pointing to the CARC and RARC lists [2] [3].
- Terminal state. Paid in full, short-paid and adjusted, written off with a reason, or transferred to collections, with the date and amount.
Payers publish 276/277 companion guides that layer their own rules on top of the X12 specification without changing it. Ask for the companion guide versions in force during the data window, because a status code that means "pended for review" at one payer can drive a different next step at another.
Why the action-code glossary must ship with the data
Collector notes are terse and coded, and without the supplier's glossary they are close to unreadable. A typical note reads like "CLD UHC, CLM IN PROC, REF# 4471, F/U 14D", and the meaningful field is often a local action code such as "RBL2" or "APL1" defined only in the billing office's configuration. Treat the glossary as part of the dataset, not documentation you can request later.
Require at minimum:
- A table of every action, status and reason code used in the window, with definition, owning team and dates the code was active.
- Mappings from local denial categories to X12 claim adjustment reason codes where the supplier maintains them [3].
- The work-queue rules that routed claims, such as aging thresholds, payer filters and balance cutoffs, since queue assignment is itself a signal.
- Notes on code retirements and renames, so a model does not treat two names for one action as different behavior.
This mirrors what made task-oriented dialogue data useful in research: the Action-Based Conversations Dataset ties agent actions to written guidelines rather than to slot values alone [5]. Your agent needs the billing office's equivalent of those guidelines.
Fields to require in an AR work-queue extract
The schema below is the minimum to request; it lets you rebuild state at every step and score outcomes. Ask suppliers to map their system exports into it, or at least to document the gaps.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value | Why it matters |
|---|---|---|
| claim_key | tokenized ID, stable across events | Joins status, actions and remits for one claim |
| event_seq / event_ts | 7 / 2025-03-14T10:22Z | Orders the trajectory; enables time-to-resolution |
| event_type | status_277, portal_check, call, action, remit_835, note | Separates observations from decisions |
| channel | clearinghouse, payer portal, phone, PM system | Agents need to know which channel was used |
| actor_role | AR rep II, denial specialist, coder, system | Distinguishes automation from human judgment |
| status_category / status_code | P1 / payer-specific code | Observation the agent conditions on |
| action_code | RBL2 (local; glossary row 18) | The decision label to predict |
| carc / group_code | 16 / CO | Denial or adjustment reason from the 835 [3] |
| note_text | redacted free text | Context the codes miss |
| next_followup_date | 2025-03-28 | Supervision for scheduling decisions |
| payer_class / plan_type | commercial, Medicare Advantage, Medicaid MCO | Policy differences across payers |
| billed / allowed / paid / adjusted | banded or relative values | Outcome scoring without exposing contracts |
| terminal_state / write_off_reason | adjusted_timely_filing | Ground truth for end-state evaluation |
| days_to_resolution | 63 | Efficiency outcome |
Two failure modes recur. First, extracts that include only notes without the code tables, which defeat next-action prediction. Second, extracts filtered to "resolved" claims, which bias the agent toward recoverable claims and hide the ones that aged out. Ask for open and written-off claims too, labeled as such.
How to evaluate a claim follow-up agent
Score agents on end state and repeated-trial reliability, not on whether one run matched a collector's note. The tau-bench work evaluates agents by comparing the final database state to a goal state and reports pass^k, the probability that all k independent trials succeed, which exposes agents that succeed once but fail inconsistently [1]. For AR, the end state is concrete: the claim reached the right terminal status with the right amount, adjustment and reason.
A practical evaluation design:
- Hold out by time and by payer. Train on earlier quarters and test on later ones, and keep at least one payer entirely out of training to measure transfer.
- Replay the observation stream. Feed the agent the 277 responses and remits in order and check each proposed action against what moved the claim forward.
- Score outcomes, not imitation. Track whether the chosen action led to payment, how many touches it took and days to resolution, alongside agreement with the human action.
- Run k trials per case. Report pass^k on a fixed case set so prompt or model changes cannot hide flakiness [1].
- Separate timely-filing risk. Weight cases near a payer's filing or appeal deadline, where a wrong wait-and-recheck action becomes an unrecoverable write-off.
Agreement with human collectors is a weak target on its own, because human queues contain wasted touches, such as repeated status calls on claims already in process. A good dataset lets you find and down-weight those.
Rights, PHI and payer-contract checks
AR histories are protected health information, and payer data inside them may be contract-confidential, so rights review covers both. Under HIPAA, the usual routes for sharing this data for training are de-identification under 45 CFR 164.514 (Safe Harbor or Expert Determination) or, more narrowly, a limited data set under a data use agreement, which is restricted to research, public health or health care operations purposes [4]. Collector notes are the hard part: free text routinely carries patient names, member IDs, callback numbers and representative names.
Questions to put to any supplier:
- Which de-identification method was used, who performed it and how dates of service were handled, since Safe Harbor removes all date elements except the year and dates are central to aging analysis [4].
- Whether the provider's business associate agreements with its billing vendor or clearinghouse permit this secondary use.
- Whether payer contracts restrict disclosure of allowed amounts or fee schedules; if so, ask for banded or relative amounts rather than raw allowed values.
- Whether payer names are kept, generalized to payer class or tokenized, and what that does to per-payer evaluation.
- Whether representative names and call reference numbers from payer calls were removed.
SourceX treats health records the same way: they require HIPAA de-identification by Safe Harbor or Expert Determination, personal details such as names, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, and no method is perfect. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Request template for AR follow-up data
Describe the data and the decision your agent makes, not the organization you think holds it. A tight request makes it easier to tell whether a supplier's extract will support training and evaluation. If you are scoping this now, you can send the request to SourceX's buyer team.
Illustrative example: invented to show structure; it does not describe an available dataset.
Use case: next-best-action agent for aged professional claims (31-180 days)
Records: claim-level event logs, first aging trigger to terminal state
Must include: 276/277 (005010X212) status pairs, portal/phone touches with
reference numbers, coded collector actions, notes, follow-up dates,
835 adjustments with group code + CARC, terminal state and amount
Glossary: full action/status/denial code tables with active dates
Coverage: mix of commercial, Medicare Advantage and Medicaid MCO payers;
include written-off and still-open claims, labeled
Window: 24+ months to allow time-based holdout
Privacy: HIPAA de-identification with method documented; amounts banded
if payer contracts restrict disclosure
Use: training and evaluation of an internal agent; terms to be agreed
Add a short note on volume expectations and the specialties you care about, such as radiology, behavioral health or durable medical equipment, since payer behavior differs sharply by specialty.
How SourceX sources AR follow-up histories
SourceX sources operational datasets from US companies on request, including support histories, engineering records, documents, and finance and legal workflows; nothing is held in stock, and a request does not guarantee a match. The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Every release is approved by the supplying company, and delivery runs through private, access-controlled workflows only after an executed agreement.
For related agent data, see SourceX's pages on enterprise workflow datasets and agent trajectories and enterprise AI agent training data, plus guides on pharmacy claim reject and resolution data and post-training data from real work.
Request AR follow-up histories for your claims agent
SourceX looks for US businesses that hold the work-queue histories you describe and manages licensing and ongoing purchases with the supplier. Each dataset is rights-reviewed and personal details are removed or replaced before delivery. Describe the AR follow-up data you need.
Sources
- arXiv (Yao et al.), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- Commonwealth of Massachusetts (MassHealth), "835 Payment Advice and EOB/CARC & RARC Lists". https://www.mass.gov/info-details/835-payment-advice
- X12, "Claim Adjustment Reason Codes". https://x12.org/codes/claim-adjustment-reason-codes
- Electronic Code of Federal Regulations, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- arXiv (Chen et al.), "Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems" (2021). https://arxiv.org/abs/2104.00783v1
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.