Skip to content

Industry-specific operational data

Health authority questions and sponsor responses for regulatory affairs AI

Quick answer

A regulatory query response dataset pairs each health authority question (an FDA information request, a deficiency, a complete response letter item, a meeting-minutes action) with the sponsor's submitted answer and what the agency did next. Public complete response letters give you the questions only [1][3]. The paired answers and outcomes sit inside sponsors' regulatory information management systems as confidential correspondence, so a usable training or evaluation set has to be licensed from the companies that hold it, with sponsor approval and confidential commercial information handled.

By SourceX Editorial · Updated

What public sources cover, and where they stop

Public sources give you agency questions without the matching sponsor answers. As of October 2026, FDA publishes complete response letters (CRLs) through openFDA: a first batch of more than 200 letters for since-approved applications appeared in July 2025 [1], and in September 2025 the agency said it would release new CRLs soon after issuance and added 89 letters tied to pending or withdrawn applications [2]. The openFDA dataset covers NDAs and BLAs and is updated infrequently [3].

That is useful retrieval context and a source of realistic deficiency language, but it is one side of the conversation. CRLs are redacted for trade secrets, confidential commercial information and personal information [2], and they do not include the sponsor's resubmission, the Type A meeting discussion or whether each deficiency was resolved. Research datasets help less than their names suggest: RegGuard, for example, builds about 967 question-answer pairs from roughly 139 regulatory documents [6], which teaches a model to read regulations, not to answer a reviewer.

The unit that actually trains a drafting assistant

The useful unit is a closed loop: question, response, and agency disposition. A drafting model trained on questions and answers alone learns to sound responsive; one trained with dispositions learns which response patterns closed the issue and which triggered a follow-up request or a second cycle.

Structure each loop at the level of the individual question, not the letter. One information request can hold a dozen questions across CMC, clinical pharmacology and labeling, and each has its own answer, attachments and outcome. HL7's Response to Regulatory Questions (RTQ) implementation guide offers a ready schema: questions and answers as FHIR Questionnaire and QuestionnaireResponse resources, linked by item identifiers and mapped to CTD sections so prior answers can be retrieved by section or meaning [4].

Sponsors already treat these archives as reusable assets. In one vendor case study, Moderna filters more than 1,600 health authority queries and answers in its RIM system to reuse earlier responses [5]. That is the shape of data to ask for: an indexed archive, not a folder of PDFs.

Record schema to request from a supplier

A clean record carries the question text, the response text, references to supporting documents, the regulatory context and the outcome as separate fields. The schema below shows the minimum a buyer should specify for RAG, supervised fine-tuning and evaluation uses.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy it matters
loop_idHAQ-2024-0311-Q07Stable key across question, answer and follow-up
authorityFDA CDERReviewer culture differs by agency and center
application_typeNDA (505(b)(2))Separates IND, NDA/BLA and device correspondence
lifecycle_stageLate-cycle, information requestEarly-stage and late-cycle questions read differently
ctd_section3.2.P.5.1Retrieval by dossier location, as in the RTQ guide [4]
disciplineCMC: specificationsRoutes to the right drafting style and reviewer
question_text"Justify the proposed acceptance criterion for impurity X..."Model input
response_textSponsor narrative with tables rendered as textModel target
attachmentsList of referenced reports, with hashesGrounding without shipping full dossiers
response_days14Signals urgency and depth
dispositionAccepted / follow-up issued / carried to CRL / post-marketing commitmentThe label that makes eval possible
follow_up_loop_idHAQ-2024-0402-Q02Links multi-round threads
redaction_logCCI masked: supplier names, batch numbersRecords what was removed and how

Two failure modes recur. First, response files often contain the final submitted letter but not the question it answers, so pairing has to be reconstructed from cover letters and tracking spreadsheets. Second, dispositions are rarely recorded explicitly; they must be inferred from the next agency communication, which needs a reviewer, not a regex.

Sponsor correspondence is confidential commercial material, and the sponsor's written approval is the gating item for any license. FDA's own rules distinguish trade secrets from confidential commercial information and give submitters notice before planned disclosure of designated material [7]; the same categories (process details, specifications, supplier identities, unpublished study results) are exactly what fills a response letter. Treat the sponsor as the rights holder for its own text and the agency's questions as context.

Where the sponsor used a CRO or regulatory consultancy to draft responses, confirm that the consultancy may release client correspondence; see the guide on client data held by service providers. Responses that quote patient-level narratives, such as safety cases cited in a clinical deficiency, carry health information. If those excerpts are protected health information, HIPAA de-identification by Expert Determination or Safe Harbor applies [8], and the Expert Determination review checklist covers what to inspect.

Reviewer and sponsor staff names also appear in signatures, meeting attendee lists and email headers. Decide up front whether you need role labels ("CMC reviewer", "regulatory lead") or nothing at all.

Matching the data to RAG, SFT and evaluation

Each application needs a different slice of the archive. Retrieval systems need breadth and accurate metadata; fine-tuning needs fewer, carefully chosen high-quality pairs; evaluation needs held-out loops with known dispositions.

Illustrative example: invented to show structure; it does not describe an available dataset.

UseWhat to licenseKey acceptance check
RAG for prior-answer reuseFull question and answer archive with CTD and product metadataRetrieval hits the right prior answer for a sample of new questions
Drafting SFTCurated pairs whose disposition was "accepted"Each response reviewed for completeness and current regulatory position
EvaluationHeld-out multi-round threads with dispositionsNo overlap with training loops by product or application
Triage or routingQuestion text with discipline and CTD labelsLabel agreement on a double-coded sample

Keep evaluation threads separate by product, not by random split. Questions on the same molecule repeat language across cycles, and a random split leaks answers. For general guidance on building pairs from business records, see turning business records into instruction-response pairs and how to source supervised fine-tuning data.

Device, safety and medical affairs correspondence are separate corpora

Device additional-information requests, pharmacovigilance queries and medical information inquiries should be treated as separate corpora, not merged with drug application correspondence. Device reviewers ask about bench testing, software documentation and predicates; drug reviewers ask about CMC, clinical pharmacology and labeling. Merging them without a product-type field teaches a model the wrong register.

Medical information responses to healthcare professionals have their own guide: medical information inquiries and standard response documents. Requirements-mapping work, which links regulations to internal controls, is covered in regulatory obligation libraries and change mapping. Browse the full industry-specific operational data hub for related correspondence types.

Questions to settle before you request data

Clarify the scope before contacting suppliers so the request describes data, not companies:

  • Agencies and centers: FDA CDER, CBER, CDRH, EMA, PMDA, others.
  • Application types and stages: IND, NDA/BLA, 505(b)(2), 510(k), De Novo, PMA; pre-submission versus review cycle.
  • Date range, given shifts in agency practice and the CRL publication policy [2].
  • Whether you need attachments, or only the narrative and references to them.
  • Required disposition labels and how inference will be documented.
  • Redaction standard for CCI and personal data, and who signs off.
  • Allowed uses: retrieval, fine-tuning, evaluation, or all three.

SourceX sources operational datasets, including documents and finance and legal workflows, from US companies on request, and manages the commercial process, including licensing agreements and ongoing purchases. Buyers can describe a regulatory correspondence request to SourceX; nothing is held in stock, and a request does not guarantee a match. More on adjacent sourcing is on the healthcare buyers page and in enterprise document datasets.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Sourcing health authority query and response data

SourceX looks for US businesses that hold the correspondence you describe, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, with personal details removed or replaced before delivery. Start a buyer request.

Sources

  1. U.S. Food and Drug Administration, "FDA Embraces Radical Transparency by Publishing Complete Response Letters" (2025). https://www.fda.gov/news-events/press-announcements/fda-embraces-radical-transparency-publishing-complete-response-letters
  2. U.S. Food and Drug Administration, "FDA Announces Real-Time Release of Complete Response Letters, Posts Previously Unpublished Batch of 89" (2025). https://www.fda.gov/news-events/press-announcements/fda-announces-real-time-release-complete-response-letters-posts-previously-unpublished-batch-89
  3. openFDA, "Complete Response Letters (CRLs)". https://open.fda.gov/apis/transparency/completeresponseletters/
  4. HL7 International, "Response to Regulatory Questions (RTQ) Implementation Guide: Use case". https://build.fhir.org/ig/HL7/rtq-ig/use-case.html
  5. Veeva Systems, "Moderna improves health authority query management with Veeva Vault RIM". https://www.veeva.com/resources/moderna-improves-health-authority-query-management-with-veeva-vault-rim/
  6. arXiv, "RegGuard: AI-Powered Retrieval-Enhanced Assistant for Pharmaceutical Regulatory Compliance" (2026). https://arxiv.org/pdf/2601.17826
  7. Legal Information Institute, "21 CFR 20.61 - Trade secrets and commercial or financial information which is privileged or confidential". https://www.law.cornell.edu/cfr/text/21/20.61
  8. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data