Industry-specific operational data
Health authority questions and sponsor responses for regulatory affairs AI
Quick answer
A regulatory query response dataset pairs each health authority question (an FDA information request, a deficiency, a complete response letter item, a meeting-minutes action) with the sponsor's submitted answer and what the agency did next. Public complete response letters give you the questions only [1][3]. The paired answers and outcomes sit inside sponsors' regulatory information management systems as confidential correspondence, so a usable training or evaluation set has to be licensed from the companies that hold it, with sponsor approval and confidential commercial information handled.
By SourceX Editorial · Updated
What public sources cover, and where they stop
Public sources give you agency questions without the matching sponsor answers. As of October 2026, FDA publishes complete response letters (CRLs) through openFDA: a first batch of more than 200 letters for since-approved applications appeared in July 2025 [1], and in September 2025 the agency said it would release new CRLs soon after issuance and added 89 letters tied to pending or withdrawn applications [2]. The openFDA dataset covers NDAs and BLAs and is updated infrequently [3].
That is useful retrieval context and a source of realistic deficiency language, but it is one side of the conversation. CRLs are redacted for trade secrets, confidential commercial information and personal information [2], and they do not include the sponsor's resubmission, the Type A meeting discussion or whether each deficiency was resolved. Research datasets help less than their names suggest: RegGuard, for example, builds about 967 question-answer pairs from roughly 139 regulatory documents [6], which teaches a model to read regulations, not to answer a reviewer.
The unit that actually trains a drafting assistant
The useful unit is a closed loop: question, response, and agency disposition. A drafting model trained on questions and answers alone learns to sound responsive; one trained with dispositions learns which response patterns closed the issue and which triggered a follow-up request or a second cycle.
Structure each loop at the level of the individual question, not the letter. One information request can hold a dozen questions across CMC, clinical pharmacology and labeling, and each has its own answer, attachments and outcome. HL7's Response to Regulatory Questions (RTQ) implementation guide offers a ready schema: questions and answers as FHIR Questionnaire and QuestionnaireResponse resources, linked by item identifiers and mapped to CTD sections so prior answers can be retrieved by section or meaning [4].
Sponsors already treat these archives as reusable assets. In one vendor case study, Moderna filters more than 1,600 health authority queries and answers in its RIM system to reuse earlier responses [5]. That is the shape of data to ask for: an indexed archive, not a folder of PDFs.
Record schema to request from a supplier
A clean record carries the question text, the response text, references to supporting documents, the regulatory context and the outcome as separate fields. The schema below shows the minimum a buyer should specify for RAG, supervised fine-tuning and evaluation uses.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value | Why it matters |
|---|---|---|
| loop_id | HAQ-2024-0311-Q07 | Stable key across question, answer and follow-up |
| authority | FDA CDER | Reviewer culture differs by agency and center |
| application_type | NDA (505(b)(2)) | Separates IND, NDA/BLA and device correspondence |
| lifecycle_stage | Late-cycle, information request | Early-stage and late-cycle questions read differently |
| ctd_section | 3.2.P.5.1 | Retrieval by dossier location, as in the RTQ guide [4] |
| discipline | CMC: specifications | Routes to the right drafting style and reviewer |
| question_text | "Justify the proposed acceptance criterion for impurity X..." | Model input |
| response_text | Sponsor narrative with tables rendered as text | Model target |
| attachments | List of referenced reports, with hashes | Grounding without shipping full dossiers |
| response_days | 14 | Signals urgency and depth |
| disposition | Accepted / follow-up issued / carried to CRL / post-marketing commitment | The label that makes eval possible |
| follow_up_loop_id | HAQ-2024-0402-Q02 | Links multi-round threads |
| redaction_log | CCI masked: supplier names, batch numbers | Records what was removed and how |
Two failure modes recur. First, response files often contain the final submitted letter but not the question it answers, so pairing has to be reconstructed from cover letters and tracking spreadsheets. Second, dispositions are rarely recorded explicitly; they must be inferred from the next agency communication, which needs a reviewer, not a regex.
Confidentiality, consent and personal data
Sponsor correspondence is confidential commercial material, and the sponsor's written approval is the gating item for any license. FDA's own rules distinguish trade secrets from confidential commercial information and give submitters notice before planned disclosure of designated material [7]; the same categories (process details, specifications, supplier identities, unpublished study results) are exactly what fills a response letter. Treat the sponsor as the rights holder for its own text and the agency's questions as context.
Where the sponsor used a CRO or regulatory consultancy to draft responses, confirm that the consultancy may release client correspondence; see the guide on client data held by service providers. Responses that quote patient-level narratives, such as safety cases cited in a clinical deficiency, carry health information. If those excerpts are protected health information, HIPAA de-identification by Expert Determination or Safe Harbor applies [8], and the Expert Determination review checklist covers what to inspect.
Reviewer and sponsor staff names also appear in signatures, meeting attendee lists and email headers. Decide up front whether you need role labels ("CMC reviewer", "regulatory lead") or nothing at all.
Matching the data to RAG, SFT and evaluation
Each application needs a different slice of the archive. Retrieval systems need breadth and accurate metadata; fine-tuning needs fewer, carefully chosen high-quality pairs; evaluation needs held-out loops with known dispositions.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Use | What to license | Key acceptance check |
|---|---|---|
| RAG for prior-answer reuse | Full question and answer archive with CTD and product metadata | Retrieval hits the right prior answer for a sample of new questions |
| Drafting SFT | Curated pairs whose disposition was "accepted" | Each response reviewed for completeness and current regulatory position |
| Evaluation | Held-out multi-round threads with dispositions | No overlap with training loops by product or application |
| Triage or routing | Question text with discipline and CTD labels | Label agreement on a double-coded sample |
Keep evaluation threads separate by product, not by random split. Questions on the same molecule repeat language across cycles, and a random split leaks answers. For general guidance on building pairs from business records, see turning business records into instruction-response pairs and how to source supervised fine-tuning data.
Device, safety and medical affairs correspondence are separate corpora
Device additional-information requests, pharmacovigilance queries and medical information inquiries should be treated as separate corpora, not merged with drug application correspondence. Device reviewers ask about bench testing, software documentation and predicates; drug reviewers ask about CMC, clinical pharmacology and labeling. Merging them without a product-type field teaches a model the wrong register.
Medical information responses to healthcare professionals have their own guide: medical information inquiries and standard response documents. Requirements-mapping work, which links regulations to internal controls, is covered in regulatory obligation libraries and change mapping. Browse the full industry-specific operational data hub for related correspondence types.
Questions to settle before you request data
Clarify the scope before contacting suppliers so the request describes data, not companies:
- Agencies and centers: FDA CDER, CBER, CDRH, EMA, PMDA, others.
- Application types and stages: IND, NDA/BLA, 505(b)(2), 510(k), De Novo, PMA; pre-submission versus review cycle.
- Date range, given shifts in agency practice and the CRL publication policy [2].
- Whether you need attachments, or only the narrative and references to them.
- Required disposition labels and how inference will be documented.
- Redaction standard for CCI and personal data, and who signs off.
- Allowed uses: retrieval, fine-tuning, evaluation, or all three.
SourceX sources operational datasets, including documents and finance and legal workflows, from US companies on request, and manages the commercial process, including licensing agreements and ongoing purchases. Buyers can describe a regulatory correspondence request to SourceX; nothing is held in stock, and a request does not guarantee a match. More on adjacent sourcing is on the healthcare buyers page and in enterprise document datasets.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Sourcing health authority query and response data
SourceX looks for US businesses that hold the correspondence you describe, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, with personal details removed or replaced before delivery. Start a buyer request.
Sources
- U.S. Food and Drug Administration, "FDA Embraces Radical Transparency by Publishing Complete Response Letters" (2025). https://www.fda.gov/news-events/press-announcements/fda-embraces-radical-transparency-publishing-complete-response-letters
- U.S. Food and Drug Administration, "FDA Announces Real-Time Release of Complete Response Letters, Posts Previously Unpublished Batch of 89" (2025). https://www.fda.gov/news-events/press-announcements/fda-announces-real-time-release-complete-response-letters-posts-previously-unpublished-batch-89
- openFDA, "Complete Response Letters (CRLs)". https://open.fda.gov/apis/transparency/completeresponseletters/
- HL7 International, "Response to Regulatory Questions (RTQ) Implementation Guide: Use case". https://build.fhir.org/ig/HL7/rtq-ig/use-case.html
- Veeva Systems, "Moderna improves health authority query management with Veeva Vault RIM". https://www.veeva.com/resources/moderna-improves-health-authority-query-management-with-veeva-vault-rim/
- arXiv, "RegGuard: AI-Powered Retrieval-Enhanced Assistant for Pharmaceutical Regulatory Compliance" (2026). https://arxiv.org/pdf/2601.17826
- Legal Information Institute, "21 CFR 20.61 - Trade secrets and commercial or financial information which is privileged or confidential". https://www.law.cornell.edu/cfr/text/21/20.61
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.