Skip to content

Text and language data

Patent Full-Text Corpora for LLM Training: Grants, Applications and Claims

Quick answer

A patent text dataset for LLM training usually starts with the USPTO's own bulk full text. Grants are published weekly from 1976 onward without images [1], and pre-grant applications generally appear 18 months after filing [7]. The work lies in parsing several legacy XML and SGML formats [3], splitting out claims, abstract and description, and linking each application to its grant. You also need to deduplicate patent families before you split train and eval. Buy packaged versions for convenience, not for rights.

By SourceX Editorial · Updated

Which patent text sources actually cover grants, applications and claims?

The authoritative source is the USPTO bulk product, and every other option is a derivative of it with different coverage. The Patent Grant Full Text product covers grants from 1976 to the present, is issued weekly, and excludes images. Drawings and some complex work units such as chemical structures and tables ship as separate files [1]. Pre-grant publication full text is a separate product. It begins with the March 2001 publications, after 35 U.S.C. 122(b) made 18-month publication the default for most applications [7].

Mirrors are the most common trap. Google's widely used USPTO bulk download page stopped updating in 2015 and points users back to the USPTO [2]. A corpus built from it silently lacks a decade of grants, including most patents that claim transformer-era software. Check the newest issue date in any third-party dump before you trust its coverage claims.

Packaged products trade cost for convenience. Snowflake's US Patents product pre-splits documents into sections such as abstract, claims and descriptions [4]. Cybersyn bundles patents with SEC filings and government contracts as public-domain material for LLM training [5]. These packages can save parsing effort. Verify field coverage before you choose one, because some database-style dumps keep bibliographic and classification fields and drop or truncate claim and description text.

Illustrative example: invented to show structure; it does not describe an available dataset.

Source typeGrantsApplicationsClaims as separate fieldImagesMain risk
USPTO grant full text (weekly)1976 onwardNoYes, after parsingSeparate productFormat changes across years
USPTO pre-grant publication full textNo2001 onwardYes, after parsingSeparate productNot all applications publish
Unmaintained mirrorPartialPartialYes, after parsingSometimes embeddedStale after the last update
Section-split warehouse productYesVariesUsuallyNoField truncation, refresh lag
Bibliographic-only database dumpYesYesOften missingNoNo claim or description text

How do you parse USPTO patent grant full-text XML without losing claims?

You need one parser per format era, a document splitter, and validation that every record has a non-empty claims block. Bulk files span a fixed-field tagged text format for the earliest years, then SGML, then several versions of XML, and community parsers exist because no single schema covers them all [3]. A weekly XML file is not one well-formed document. It is many complete documents concatenated, each with its own XML declaration and DOCTYPE, so a naive etree.parse fails or reads only the first patent.

For the modern XML grant format, the elements you will map most often are:

  • us-bibliographic-data-grant, holding publication-reference (document number, kind code, date) and application-reference.
  • classifications-cpc and the IPC fields, which you need for domain filtering and stratified evaluation.
  • abstract, description (with heading elements such as background and summary) and claims.
  • Each claim with an id and num, nested claim-text, and claim-ref idref linking dependent claims to their parents.

The common failure modes are concrete. Nested claim-text elements flatten into run-on strings unless you preserve indentation or separators. Inline tables, maths and chemistry elements point to external files, and you need a placeholder policy. Entity references such as ‘ break if you parse without the DTD. Character encoding and entity conventions differ between the SGML and XML eras, so normalize everything to UTF-8 before deduplication.

What should a training-ready patent record contain?

A useful record keeps claims structured, retains the identifiers that link applications to grants, and records which parser version produced it. Patent documents are long and terminologically dense [6], so section boundaries matter for chunking, for long-context training, and for any task that conditions on one section to produce another. The schema below follows the fields recommended in our guide to metadata fields for licensed text corpora.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "doc_id": "US-0000000-B2",
  "kind_code": "B2",
  "doc_type": "utility_grant",
  "application_number": "00/000,000",
  "earliest_priority_date": "2019-03-14",
  "publication_date": "2023-06-20",
  "related_pgpub_id": "US-2020-0000000-A1",
  "family_id": "FAM-000000",
  "cpc_main": "G06F 40/30",
  "title": "System for routing support tickets",
  "abstract": "...",
  "description_sections": [{"heading": "BACKGROUND", "text": "..."}],
  "claims": [
    {"num": 1, "type": "independent", "category": "method", "depends_on": null, "text": "..."},
    {"num": 2, "type": "dependent", "depends_on": 1, "text": "..."}
  ],
  "source_file": "weekly grant file, issue date 2023-06-20",
  "parser_version": "grant-xml-v4.x parser 1.3.0",
  "excluded_elements": ["drawings", "chemistry", "tables"]
}

Kind codes deserve their own field. A1 and A2 mark pre-grant publications, B1 and B2 mark utility grants, and S, P and E mark design, plant and reissue documents. Design patents carry almost no text beyond a single claim, so filter them out of claim-generation data or they will teach the model one-line boilerplate.

How do patent families and publications cause duplication and eval leakage?

The same invention appears several times in the USPTO stream, so random splits leak. An application publishes as an A1 document and later grants as a B2 with largely the same description; B1 marks grants that were never published as applications. Continuations and divisionals repeat the parent specification almost verbatim. A model evaluated on a random held-out sample has usually already seen most of each test description.

Use these controls before splitting:

  1. Group by family or by application number plus continuity data, and assign whole groups to a single split.
  2. Run near-duplicate detection on descriptions, because continuation text overlaps even when the identifiers differ. See exact and near-duplicate detection with MinHash and LSH.
  3. Split evaluation by priority date for generation tasks, so the test set postdates the training text.
  4. Check overlap with public web crawls. Patent text is widely mirrored and probably already sits in general pre-training mixes; see testing licensed data against public web corpora.

The A1-to-B2 pair is also a signal, not only a nuisance. Claims are often amended during prosecution, so pairing published-application claims with granted claims gives you aligned before-and-after text. The reasons for each change sit in the file wrapper (office actions and responses), not in the full-text product.

Which tasks does a public patent corpus support, and where does it stop?

Public full text supports continued pre-training, claim generation, classification and prior-art retrieval, but it does not show how practitioners draft, argue or decide. Researchers have built claim-generation datasets from European patents to go beyond USPTO-only data [8]. Others pair generated patent text with prior-art search and reranking [9], and some train patent models with instruction tuning and human feedback [10]. Each of these depends on clean section and claim boundaries.

For continued pre-training on domain corpora, weight descriptions and claims differently. Descriptions carry technical vocabulary. Claims carry a narrow legal register with antecedent-basis rules ("a sensor ... the sensor") that models reproduce badly when they see claims only as flattened text. For retrieval, index claims and description paragraphs as separate units with CPC metadata, and evaluate against citation data rather than keyword overlap.

The full-text record stops at the publication. Office actions and applicant responses are public for published applications but sit in separate prosecution records, while invention disclosures, internal drafts, claim charts and the drafting reasoning behind amendments are never published. Our guide to patent drafting and prosecution data covers those non-public artifacts and how buyers source them.

What rights and documentation questions still apply to public-domain patent text?

USPTO patent text is generally treated as public domain, and the grant full-text catalog entry carries a public-domain mark [1], but your documentation duties still apply. A small number of specifications include third-party material or copyright notices, and drawings and embedded images follow separate products [1]. Keep a filter list for such documents rather than assuming every byte is unencumbered. Foreign patent office data, including European sources [8], comes under each office's own terms of use, so do not apply USPTO assumptions to EPO, JPO or CNIPA text.

Disclosure rules apply even to public-domain inputs. A general-purpose AI model provider in the EU must keep a copyright policy and publish a training-content summary under Article 53 of the AI Act [11]. As of October 2026, those duties have applied since 2 August 2025. California AB 2013 requires generative AI developers to post documentation about their training datasets [12]. Record the source product, issue-date range, parser version and filters for each patent shard so that these summaries are accurate.

Packaged vendor products add a contract layer. A warehouse listing may be built from public-domain text and still restrict redistribution, derived datasets or use outside the platform. Read the pre-training license rights you are actually being granted, and compare it with the free USPTO source before you pay for convenience.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Patent corpus build checklist

Run this checklist before a patent corpus enters a training mix.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Source is the USPTO bulk product or a derivative whose newest issue date you have verified [1][2].
  • Grants and pre-grant publications are both included if your task needs applications.
  • Parser handles each format era; record count per weekly file matches the file's document count [3].
  • Every utility record has non-empty claims; dependent claims resolve to a parent claim-ref.
  • Design, plant and reissue documents are tagged by kind code and filtered per task.
  • Tables, math, chemistry and drawings have an explicit placeholder or exclusion policy.
  • Family-level grouping and near-duplicate removal run before the train/eval split.
  • Eval split is date-based for generation tasks; overlap with web crawls is measured.
  • Vendor package terms are checked for redistribution and derived-data limits [4][5].
  • Training-data documentation records sources, date ranges and filters [11][12].

For broader context on the cluster, start at text datasets for LLM training, or compare adjacent public corpora such as SEC filings text and open-access scientific full text.

Sourcing patent and IP work data beyond public full text

Public patent text is free to download, so most teams do not need a sourcing partner for it. The data that differentiates patent and legal models is the work around publication: documents, legal workflows and engineering records held inside companies. That is what SourceX's buyer program addresses. Use cases appear on our legal AI training data page.

Request non-public patent and IP work data

SourceX sources operational datasets from US companies on request, including documents, engineering records and legal workflows. Nothing is held in stock and a request does not guarantee a match. Each dataset is rights-reviewed, approved by the supplying company and delivered under a license that defines records, uses, term and delivery. Describe the data you need at https://sourcex.si/buyers.

Sources

  1. Data.gov (U.S. Patent and Trademark Office dataset), "Patent Grant Full Text (1976 - Present)". https://catalog.data.gov/dataset/patent-grant-full-text-1976-present
  2. Google, "Google USPTO Bulk Downloads: Patent Grant Full Text with Embedded Images". https://www.google.com/googlebooks/uspto-patents-redbook.html
  3. GitHub (TamerKhraisha), "uspto-patent-data-parser". https://github.com/TamerKhraisha/uspto-patent-data-parser
  4. Snowflake (data documentation), "US Patents". https://data-docs.snowflake.com/foundations/products/us-patents
  5. Cybersyn documentation, "LLM Training (public domain technology data)". https://docs.cybersyn.com/public-domain/technology/llm-training
  6. arXiv (2403.04105), "Natural Language Processing in the Patent Domain: A Survey" (2024). https://arxiv.org/pdf/2403.04105
  7. BitLaw, "35 U.S.C. 122: Confidential status of applications; publication of patent applications". https://bitlaw.com/source/35usc/122.html
  8. arXiv (2505.12568), "Enriching Patent Claim Generation with European Patent Dataset" (2025). https://arxiv.org/pdf/2505.12568
  9. arXiv (2009.09132), "Prior Art Search and Reranking for Generated Patent Text" (2020). https://arxiv.org/pdf/2009.09132
  10. arXiv (2406.16897), "InstructPatentGPT: Training patent language models to follow instructions with human feedback" (2024). https://arxiv.org/pdf/2406.16897
  11. EUR-Lex, "Regulation (EU) 2024/1689 (AI Act, consolidated)" (2024). https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:02024R1689-20260727
  12. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data