Industry-specific operational data
FNOL intake conversations for voice and chat claims-intake agents
Quick answer
FNOL data for AI is a set of first-notice-of-loss conversations (telephony audio or chat logs) joined to the structured loss notice the handler created and to what happened next: loss cause, severity or complexity tier, assignment queue and later flags such as litigation or fraud. Buyers training voice or chat intake agents need all three layers, recorded with documented consent, with claimant and third-party identifiers and injury details redacted, and licensed from the carriers, MGAs or TPAs that hold them.
By SourceX Editorial · Updated
What a usable FNOL record contains
A usable FNOL record pairs the dialogue with the notice it produced and the triage decision that followed. The conversation alone teaches turn-taking; the notice teaches slot filling; the triage outcome teaches routing. Without the join keys between them, you have a speech corpus, not intake training data.
The notice layer usually mirrors the fields of the industry's standard loss-notice forms, such as the ACORD property, automobile and general liability notices, which many carriers use as a reference schema. Expect date, time and location of loss, cause description, policy number and coverage verification result, involved parties and vehicles, injury indicator, police or fire report number, and reporter relationship (insured, claimant, agent, attorney). Ask suppliers which fields were captured live during the call versus completed later by an adjuster, because only live-captured fields are valid targets for an intake agent.
The outcome layer is where most of the value sits. Useful labels include loss-cause code, coverage line, severity or complexity tier, fast-track versus field assignment, total-loss flag, subrogation potential, and indicators set weeks later, such as attorney representation or SIU referral. Those late labels are powerful for triage models and also carry the historical bias described in auditing outcomes in operational labels.
Voice versus chat FNOL: what changes for the buyer
Voice and chat FNOL data differ in signal quality, consent mechanics and preparation cost, so specify which one your model actually consumes. Many intake vendors need both, but in different proportions.
Telephony FNOL is typically 8 kHz narrowband audio, often G.711 mu-law, sometimes stored as mixed mono and sometimes as dual-channel stereo with agent and caller on separate channels. Dual-channel recordings make diarization and per-speaker redaction far easier, so make channel separation a hard requirement if your ASR or speech-to-speech model was trained on wideband audio. Callers are frequently distressed, on the roadside, multilingual or calling through an interpreter line, which is exactly the condition public corpora lack: Common Voice, for example, is crowdsourced read speech released into the public domain, not stressed telephony dialogue [8]. Recent research on call-center speech also points to how few such corpora are openly available and how often their terms limit commercial use [7].
Chat and web-form FNOL arrive as structured JSON or platform exports with timestamps, bot turns, handoff events and uploaded photo references. They are cheaper to redact and easier to convert into the role-tagged format covered in chat fine-tuning data format, but they under-represent injury and multi-vehicle losses, which tend to arrive by phone.
Consent, biometrics and health details in FNOL recordings
FNOL audio carries three distinct legal exposures: call-recording consent, voiceprint biometrics and health information about injured parties. Each needs its own evidence in diligence.
Recording consent is set by state law and covers everyone on the call, not just the policyholder. California makes it an offense to record a confidential communication without the consent of all parties [1], and section 632.7 separately covers calls involving cellular or cordless phones, which describes most roadside FNOL calls [2]. Ask for the IVR disclosure script, the dates it was in use, and how third-party claimants and conferenced parties were notified.
Voice is also a biometric question. Texas treats voiceprints as biometric identifiers that require notice and consent before capture for a commercial purpose [4], and class actions filed in May 2026 allege that voiceprints used to train AI fall under Illinois' BIPA; as of October 2026 those claims have not been decided on the merits [3]. If you will train speaker-dependent or speech-to-speech models, ask whether the supplier's notices covered that purpose and whether voice conversion or speaker anonymization is an option.
Injury descriptions in auto and liability FNOL are health details even when the carrier is not a HIPAA covered entity. Washington's My Health My Data Act defines consumer health data broadly, including biometric data, although it exempts some information governed by other laws such as the Gramm-Leach-Bliley Act, so confirm with counsel whether it reaches your supplier's records [6]. Where records do come from a covered entity's workflow, HIPAA de-identification under Safe Harbor or Expert Determination applies, and HHS notes that neither method removes all re-identification risk [5]. This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Redaction and preparation failure modes to test
Redaction in FNOL data fails in predictable places, so test those places in the sample rather than accepting a method description. Text redaction on the transcript does not remove the same spans from the audio, so ask for time-aligned audio bleeping or replacement.
Common failures to check:
- Policy, claim and VIN numbers read digit by digit, which ASR splits across tokens and pattern-based scrubbers miss.
- License plates, street addresses and intersections inside the loss description, which is also the field your cause classifier needs.
- Names of third parties, witnesses, tow yards and body shops, which are identifiers for people other than the caller.
- Callback numbers in hold-queue or voicemail segments outside the main dialogue.
- Over-redaction that replaces every location with one token, destroying the geography a severity model may rely on; ask for typed surrogates such as [CITY] or [INTERSECTION].
Then check label provenance. Severity tiers are often re-scored by adjusters days after intake, so confirm whether the label reflects what was knowable during the call. For agentic intake, the follow-on decisions belong to a different dataset; see claims adjudication decisions with coverage reasoning.
Illustrative FNOL record and request template
The fastest way to align with a supplier is to show the record shape you expect and the minimum checks you need. The example below is a schema to adapt, not a sample of real data.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"fnol_id": "F-000123",
"channel": "voice",
"audio": {"codec": "pcm_mulaw", "sample_rate_hz": 8000, "channels": 2, "redaction": "time_aligned_tone"},
"transcript_turns": [
{"speaker": "agent", "start_s": 0.0, "text": "Thank you for calling claims. This call is recorded."},
{"speaker": "caller", "start_s": 4.2, "text": "I was rear-ended at [INTERSECTION] about an hour ago."}
],
"loss_notice": {
"line": "personal_auto", "loss_date": "[DATE]", "loss_location_state": "TX",
"cause_text": "rear-end collision at stop", "police_report": true,
"injury_reported": true, "coverage_verified": true, "fields_captured_live": ["cause_text", "injury_reported"]
},
"triage": {"cause_code": "AUTO_REAR_END", "complexity_tier": 2, "queue": "bodily_injury", "set_at": "intake"},
"later_flags": {"attorney_rep": false, "siu_referral": false, "observed_after_days": 90},
"consent": {"disclosure_script_version": "v3", "two_party_state_call": false}
}
Illustrative example: invented to show structure; it does not describe an available dataset.
| Request field | What to specify |
|---|---|
| Lines of business | Personal auto, homeowners, commercial property, GL, workers' comp |
| Channels and mix | Voice, chat, web form; target share of each |
| Audio requirements | Sample rate, dual-channel, codec, minimum call length |
| Languages | English plus named languages, interpreter-line handling |
| Join keys | Conversation to notice to triage label to later flags |
| Label timing | Labels set at intake versus re-scored later |
| Consent evidence | Disclosure script, two-party-state handling, third-party notice |
| Redaction | Text and audio, typed surrogates, sample QA |
| Intended use | Fine-tuning, evaluation, classification, summarization |
If you are building held-out test sets rather than training data, the scenario and outcome design in voice agent evaluation sets and insurance claims AI evaluation apply directly.
How this differs from whole-claim files and generic call audio
This page covers the intake conversation and its immediate triage, not the full claim lifecycle and not generic contact-center audio. Whole claim files, including adjuster notes, reserves and payments, are covered on insurance claims datasets for AI training. Broader speech needs, such as accent coverage across many industries, sit with call center audio datasets and voice agent training data.
Post-intake servicing calls, such as status checks and escalations, show different intents and are discussed in insurance service escalation records. For the wider set of insurance buyer needs, start from buyers by industry: insurance or the industry-specific operational data hub.
How SourceX handles an FNOL data request
SourceX sources operational datasets, including support histories and other operational records, from US companies on request; it does not hold FNOL data in stock, and a request does not guarantee a match. You describe the conversations, notices and labels you need, and SourceX looks for US businesses that hold them; every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details such as names, phone numbers and account numbers are removed or replaced with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows after an executed agreement. You can submit a buyer request with the template fields above.
Sourcing FNOL data for AI intake agents
If your intake agent needs real first-notice-of-loss conversations joined to loss notices and triage outcomes, SourceX can look for US carriers, MGAs or service providers that hold them and manage the licensing process. Nothing is contracted until a supplier agrees, and terms are set per deal. Describe the FNOL data you need.
Sources
- California Legislative Information, "California Penal Code section 632 (recording confidential communications)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN§ionNum=632
- California Legislative Information, "California Penal Code section 632.7 (cellular and cordless telephone communications)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN§ionNum=632.7
- Biometric Update, "Tech giants sued under BIPA over voiceprints used to train AI" (2026). https://www.biometricupdate.com/202605/tech-giants-sued-under-bipa-over-voiceprints-used-to-train-ai
- Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier" (2026). https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Washington State Legislature, "Chapter 19.373 RCW - Washington My Health My Data Act". https://app.leg.wa.gov/RCW/default.aspx?cite=19.373&full=true
- arXiv, "arXiv preprint 2507.02958 (call-center speech data)" (2025). https://arxiv.org/abs/2507.02958
- Ardila et al. (Mozilla), arXiv, "Common Voice: A Massively-Multilingual Speech Corpus" (2019). https://arxiv.org/abs/1912.06670v1
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.