Text and language data
Patent Drafting and Prosecution Data: Invention Disclosures, Drafts and Office Action Responses
Quick answer
Patent drafting AI training data is the private work product that sits between an inventor's idea and a published patent: invention disclosure forms, attorney draft claims and specifications with revision history, examiner rejections paired with the response arguments and claim amendments that overcame them, and the final allowed text. Public grants show only the end state. Buyers license the drafting pipeline from IP practices and corporate patent departments, after privilege, client-confidentiality and pending-application checks.
By SourceX Editorial · Updated
What public patent text is missing for drafting models
Published patents and applications show the finished document, not the decisions that produced it. A model trained only on grant full text learns what allowed claims look like, but not how a practitioner turned a two-page disclosure into an independent claim, which limitations were added after a 35 U.S.C. 103 rejection, or why an argument was chosen over an amendment. If your goal is bulk published text, the patent full-text corpus guide covers grant and application feeds; this page is about the private layer.
Research supports the gap. Instruction-following patent models are trained with human feedback on drafting tasks [1], claim-generation work depends on purpose-built patent datasets [2], and surveys note that patent drafting is hard because of specialized terminology and very long documents [3]. Each of these points to expert-labeled pairs, not more raw text.
Which drafting and prosecution artifacts are worth licensing
The most useful records are linked pairs that show an input, an expert transformation and an outcome. A disclosure on its own has limited value; a disclosure joined to the filed claims and the attorney's redline history becomes supervised fine-tuning (SFT) and preference data. The core artifact types are:
- Invention disclosure forms (IDFs): inventor narrative, problem statement, embodiments, known prior art, figures.
- Draft claim sets and specifications: successive versions with tracked changes, partner review comments and the filed version.
- Office action packages: the rejection, cited references, the response (remarks plus amended claims), interview summaries and the next action.
- Internal strategy notes: claim-amendment rationales, continuation decisions and abandonment reasons, which carry the heaviest privilege load.
Public prosecution data is a useful complement. Published file histories show examiner rejections under 35 U.S.C. 102 and 103, cited references and the applicant's filed responses, which give you examiner-side structure; private files supply the drafts, internal reasoning and rejected alternatives behind each response. Verify official USPTO access routes and terms before combining the two.
How prosecution histories map to SFT, preference and eval data
Each artifact type maps to a specific training format, and the mapping should drive what you request. A disclosure-to-claims pair is an SFT example. A draft claim and its partner-revised version form a natural chosen/rejected preference pair, the same structure used in human-feedback patent training [1]. An office action plus the response that led to allowance is a held-out evaluation item, because you can score a model's response against a known outcome.
Volume matters less than curation for the SFT layer. LIMA showed strong instruction following from about 1,000 carefully selected prompt-response pairs [4], so a few thousand well-linked drafting episodes from experienced practitioners can be worth more than a large unlinked dump. For post-training planning across domains, see data sourcing for post-training teams.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value | Use |
|---|---|---|
episode_id | PD-000412 | Join key across artifacts |
tech_center | 2100 (computer architecture and software) | Domain stratification |
disclosure_text | Inventor narrative, de-identified | SFT input |
claims_v1 / claims_filed | Associate draft / filed set | Preference pair (rejected / chosen) |
oa_type | non-final; 103 over two references | Eval conditioning |
response_remarks | Arguments, de-identified | SFT target |
amendment_diff | Added limitation to claim 1 | Reasoning label |
outcome | allowed after RCE | Eval ground truth |
publication_status | published | Rights gate |
Confidentiality, privilege and pending-application limits
Three legal constraints shape what can be licensed, and all three should be checked before any file moves. First, under 35 U.S.C. 122 the USPTO keeps applications in confidence until publication, which for most applications comes about 18 months after the earliest priority date, and applicants who are not filing abroad can request nonpublication. Unpublished and nonpublication-requested applications should be excluded or held until publication, because the client's trade secrets are still in them.
Second, attorney-client communications and work product belong to the client relationship, not to the firm alone. A firm usually needs client consent or an engagement-letter basis to release disclosures and strategy notes, and counsel must assess waiver risk; our insight on whether law firms can license data without waiving privilege covers this in depth. The provenance guide on client data held by service providers sets out the authorization chain buyers should ask for.
Third, inventor names, addresses, employee IDs and client matter numbers need removal or replacement. Technical content can still identify an unpublished invention, so de-identification does not fix a confidentiality problem; only exclusion or publication does.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Buyer request checklist for patent drafting data
A precise request filters out unusable supply early. Use this checklist when you describe the data to any supplier:
- Scope: technology areas, USPTO technology centers or art units, filing years, utility vs. design, US only or with PCT/EP counterparts.
- Linkage: require a shared episode key across disclosure, drafts, actions and responses; reject unlinked document dumps.
- Publication gate: only published or granted matters, with the publication number recorded per episode.
- Client authorization: written basis for each client's matters, and a list of excluded clients.
- Version history: native DOCX with tracked changes, or extracted diffs, so drafts are not flattened to the final version.
- Practitioner metadata: role (agent, associate, partner) and years of experience, de-identified, to weight preference labels.
- Format: JSONL per episode with text fields, plus original PDFs for office actions; figures as separate files.
- Contamination control: flag episodes whose final text appears in public grant corpora you already train on, so eval splits stay clean.
Related reading: legal LLM fine-tuning data covers redlines and privilege across practice areas, and proprietary text beyond web crawls explains why this kind of work product is absent from public corpora.
How SourceX approaches patent drafting data requests
SourceX sources operational datasets from US companies on request; nothing is held in stock, and a request does not guarantee a match. Legal workflows and documents are among the kinds of data it looks for, and you describe the data you need rather than naming specific firms. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, allowed uses, term and delivery, with each release approved by the supplying company.
Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement. You can describe your patent drafting data needs to SourceX, and see the broader legal AI training data use case and legal briefs and memos pages. More text categories sit under the text and language data hub and the AI data hub.
Request patent drafting and prosecution training data
If your drafting, office action response or claim-generation model needs linked disclosures, drafts and responses that public patents lack, describe the records, fields and allowed uses you need. SourceX looks for US businesses that hold the described data and manages the Find, Assess, Agree, Transact and Manage process; nothing is contracted until a supplier agrees. Start a buyer request at SourceX.
Sources
- arXiv, "InstructPatentGPT: Training patent language models to follow instructions with human feedback" (2024). https://arxiv.org/pdf/2406.16897
- arXiv, "Enriching Patent Claim Generation with European Patent Dataset" (2025). https://arxiv.org/pdf/2505.12568
- arXiv, "Natural Language Processing in the Patent Domain: A Survey" (2024). https://arxiv.org/pdf/2403.04105
- arXiv, "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.