Skip to content

Industry-specific operational data

Pharmacovigilance case narratives and ICSRs for safety case-processing AI

Quick answer

A pharmacovigilance case narrative dataset worth licensing pairs de-identified individual case safety reports (ICSRs) in ICH E2B(R3) structure with the source documents they were built from, MedDRA-coded events, the final narrative, and the safety team's seriousness, causality and expectedness assessments, ideally with follow-up versions and QC findings. Public resources such as FDA FAERS extracts and the PHEE corpus give structured fields or literature annotations, not the source-to-case pairs that intake, coding and narrative-generation models need [3][6].

By SourceX Editorial · Updated

What public adverse event data leaves out

Public adverse event data gives you coded counts, not the case-processing work itself. FAERS quarterly extracts are raw, non-cumulative files in ASCII or XML organized around drug, reaction, outcome and report-source tables, and they are meant for relational analysis rather than document AI [3]. FDA's dashboard notes that a report does not establish that the drug caused the event, so FAERS rows carry no company causality label you can train on [4]. As of October 2026, FDA's file page also references a consolidated adverse event monitoring platform, so confirm file locations before building a pipeline on them [3].

The academic picture is thin. PHEE, a prominent public pharmacovigilance text corpus, annotates more than 5,000 adverse and potential therapeutic events from published case reports and literature, and its authors note how few public PV text resources exist [6]. That makes it useful for event-extraction pretraining or a sanity-check eval, but it contains no MedWatch forms, CIOMS I forms, call-center transcripts, or company narratives written to a sponsor's conventions.

The ICSR record a model actually learns from

A trainable case is a linked bundle of inputs, structured outputs and human decisions. The E2B(R3) data-element tables define the target structure: sender and receiver headers, primary source and reporter, patient characteristics, reactions/events, test results, drug information, and a narrative section (H) that also holds reporter comments, sender diagnoses and sender comments [2]. FDA's implementation guide, Revision 1, adds US regional elements through a separate technical specification, so ask whether exports follow ICH-only or FDA-regional validation [1]. Narrative numbering changed between versions (B.5.1 in R2 versus the H section in R3); verify element IDs against the current ICH guide before you map fields [2].

Illustrative example: invented to show structure; it does not describe an available dataset.

case_bundle:
  case_id: "PSEUDO-7F3A91"          # tokenized; original safety-database ID never shipped
  version: 3                         # initial + two follow-ups, each kept
  report_type: spontaneous           # or study, literature, other
  source_documents:
    - {type: medwatch_3500, format: pdf, pages: 4, redaction: applied}
    - {type: call_center_transcript, format: txt, redaction: applied}
    - {type: discharge_summary, format: pdf, ocr: true}
  e2b_r3_export: case_v3.xml         # validated against ICH schema; regional rules noted
  meddra:
    version: "record version used"
    events: [{verbatim: "felt dizzy and passed out", llt: "<code>", pt: "<code>"}]
  assessments:
    seriousness: [hospitalization]
    causality: {reporter: possible, company: possible, method: "WHO-UMC"}
    expectedness: {reference: "CCDS/IB version", result: unexpected}
  narrative: {text: narrative_v3.txt, template: "sponsor-specific"}
  qc_findings: [{field: "onset_date", issue: "transposed", corrected_in: 2}]

Every block above maps to a model task. Source documents plus the E2B export train intake and field extraction; verbatim-to-LLT pairs train auto-coding; source-plus-structured-fields to final narrative is the supervised pair for narrative-generation SFT; and QC findings plus version diffs give you error labels and an eval set that measures what reviewers actually correct.

MedDRA licensing and coding labels

Commercial use of MedDRA terms generally requires a subscription, so check your own license before you treat coded fields as labels. The terminology is maintained by the MSSO, ICH holds the intellectual property, and commercial organizations subscribe at revenue-tiered rates [5]. A dataset delivered with LLT and PT codes does not grant you MedDRA rights; your team needs its own subscription, and the supplier's license should state which MedDRA version each case was coded in.

Version drift is the main coding failure mode. A case coded in an older version may map to a different PT after an upgrade, and a model trained across mixed versions learns inconsistent targets. Ask for the version per case, the recoding policy the supplier followed at upgrades, and whether verbatim reporter terms were preserved alongside the codes.

Seriousness, causality and expectedness as labels

Assessment fields are the most valuable and the noisiest labels in a safety database. Seriousness criteria (death, life-threatening, hospitalization, disability, congenital anomaly, other medically important) are fairly reproducible. Causality is not: reporter and company assessments often differ, methods vary (WHO-UMC categories, Naranjo scores, binary related/not related), and expectedness depends on which reference safety information version applied on the case date.

For training, keep reporter and company causality as separate fields, record the method, and carry the reference-document version for expectedness. If a supplier can only give you the final company assessment, treat the dataset as suitable for extraction and narrative work but weak for causality modeling.

De-identifying cases, narratives and source documents

De-identification has to cover the narrative and the scanned sources, not just the structured fields. ICSRs carry patient initials, birth dates, ages, event and hospitalization dates, reporter names and contact details, and site identifiers, and narratives routinely repeat them in free text. Under HIPAA, OCR describes two routes: Safe Harbor, which removes 18 identifier types including those of relatives and household members, and Expert Determination [7][8]. Safe Harbor's date rules strip the onset-to-event intervals that narratives depend on, which is why PV datasets often rely on Expert Determination with date shifting; see what dates and ZIP codes lose under Safe Harbor.

A limited data set under 164.514(e) keeps some dates but requires a data use agreement and is not de-identified data [8]. NIST also cautions that traditional de-identification has inherent limits, so ask for residual-risk reasoning, not just a redaction list [9]. Reporter identities matter separately: healthcare professional reporters are not patients, but their names and institutions still need removal. Human redaction decisions on these documents are themselves a training asset; see redaction training data.

Buyer due-diligence checklist for ICSR datasets

Ask these questions before any sample review; each one maps to a known failure in PV model projects.

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckWhat to ask forFailure it prevents
Source-to-case linkageSource document IDs linked to each case versionTraining on outputs with no inputs
E2B conformanceICH vs regional schema, validation logSilent field misalignment
MedDRA versionVersion per case, recoding policy, verbatims keptInconsistent coding targets
Assessment provenanceReporter vs company causality, method, RSI versionCollapsed or ambiguous labels
Follow-ups and QCAll versions plus QC finding codesNo error signal for evals
Narrative conventionsTemplates used, number of authoring teamsOverfitting to one style
De-identificationMethod, date handling, residual-risk notes, sample auditRe-identification from free text
Case mixSpontaneous, study, literature; product classesNarrow generalization
DocumentationA data card covering sources, preparation, intended use [10]Undocumented provenance at audit

Narrative style diversity deserves weight. Each safety organization writes narratives to its own template and house phrasing, so a model trained on one company's cases tends to reproduce that template. Related operational records with the same structure-plus-narrative shape include clinical study reports for medical writing AI and report-generation fine-tuning data.

How SourceX handles requests for safety case data

SourceX sources operational datasets from US companies on request and manages the commercial process, including the licensing agreement. Nothing is held in stock, and a request does not guarantee a match: buyers describe the data they need, SourceX looks for US businesses that hold it, and every release is approved by the supplying company. You can describe your ICSR requirements on the buyers page.

Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Health records require HIPAA de-identification by Safe Harbor or Expert Determination, and personal details are removed or replaced before delivery, with the method recorded and a sample checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. For broader context, see the industry data hub, healthcare buyers, whether AI labs buy medical data and licensing medical records.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Requesting pharmacovigilance case narrative data

Describe the case types, E2B fields, source documents, assessment labels and volumes your intake, coding or narrative models need. SourceX will look for US companies that hold matching data, review rights and de-identification, and agree allowed uses in a license before anything is delivered. Start a pharmacovigilance data request.

Sources

  1. U.S. Food and Drug Administration, "E2B(R3) Electronic Transmission of Individual Case Safety Reports Implementation Guide: Data Elements and Message Specification" (2022). https://www.fda.gov/regulatory-information/search-fda-guidance-documents/electronic-transmission-individual-case-safety-reports-implementation-guide-data-elements-and
  2. ICH (copy hosted by PMDA), "ICH E2B(R3) Data Elements for Transmission of Individual Case Safety Reports, Implementation Guide, Module III". https://www.pmda.go.jp/files/000279433.pdf
  3. U.S. Food and Drug Administration, "FDA Adverse Event Reporting System (FAERS): Latest Quarterly Data Files". https://www.fda.gov/drugs/fdas-adverse-event-reporting-system-faers/fda-adverse-event-reporting-system-faers-latest-quarterly-data-files
  4. U.S. Food and Drug Administration, "FDA Adverse Event Reporting System (FAERS) Public Dashboard FAQ". https://fis.fda.gov/extensions/FPD-FAQ/FPD-FAQ.html
  5. MedDRA MSSO, "Subscription Rates". https://www.meddra.org/subscription-rates
  6. Sun et al., "PHEE: A Dataset for Pharmacovigilance Event Extraction from Text" (2022). https://arxiv.org/pdf/2210.12560
  7. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  8. eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  9. National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
  10. Pushkarna, Zaldivar, Kjartansson, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data