Skip to content

Text and language data

SEC Filings Text Corpora for Financial LLMs

Quick answer

SEC filings are one of the largest freely accessible sources of formal financial prose: 10-K and 10-Q narratives, proxy statements (DEF 14A) and exhibits, all public on EDGAR. Usable corpora already exist, from EDGAR-CORPUS's itemized 10-Ks for 1993 to 2020 [1] to multi-billion-token cleaned releases [2]. What decides training value is preprocessing: which sections you keep, how you handle inline XBRL tables, how hard you deduplicate boilerplate, what share of the mixture filings get, and whether filings-based benchmarks stay out of training [3][4][6].

By SourceX Editorial · Updated

What existing EDGAR corpora give you, and what they leave out

Public EDGAR corpora give you volume and structure but no labels, no proprietary context and no guarantee of a clean evaluation split. EDGAR-CORPUS is the reference point: every US 10-K from 1993 to 2020, each report split into items (Item 1 Business, Item 1A Risk Factors, Item 7 MD&A and so on) and stored as JSON, without annotations [1]. That item split is its main value, because it lets you filter by section instead of training on whole documents.

Newer releases quote larger "clean token" counts; one vendor case study reports about 43B clean tokens from EDGAR, with 10-Ks the largest share [2]. Treat such figures as self-reported and tokenizer-dependent: ask which tokenizer, which form types, which date range and what was removed before the count was taken. Two releases with the same headline count can differ sharply in how much table residue, exhibit boilerplate and duplicate text they keep.

What none of these corpora contain is the operational record behind the filing: general ledgers, close checklists, reconciliation notes, auditor correspondence, credit memos. Filings are the polished output of a finance function, not the workflow that produced it. If your target task is an accounting or analyst agent, filings teach the register and vocabulary but not the procedures (see finance and accounting AI training data).

Which form types and sections to keep for continued pretraining

Keep the narrative sections and drop most of the rest. A recent scaling study built a 400M-token continued-pretraining corpus from 10-K, 10-Q and DEF 14A filings and kept narrative sections such as MD&A and Risk Factors rather than full documents [3]. Narrative sections carry the reasoning-like text (causal explanations of revenue changes, liquidity discussion, risk disclosure) that general web crawls underrepresent.

The low-value material is predictable. Cover pages, signature blocks, exhibit indexes, certifications under SOX sections 302 and 906, and the forward-looking-statement safe-harbor paragraph repeat near-verbatim across thousands of filings. Financial statement tables flattened to text become long runs of numbers and whitespace that teach the model little and inflate token counts.

Form choice should follow the use case:

  • 10-K for annual narrative depth, Risk Factors and MD&A.
  • 10-Q for higher-frequency updates and quarter-over-quarter language, with more repetition from the prior quarter.
  • DEF 14A for governance, executive compensation and shareholder-proposal language.
  • 8-K for event-driven disclosure (earnings releases are often attached as Exhibit 99.1); short documents with heavy template reuse.

How to handle inline XBRL and tables during extraction

Decide explicitly whether tables become text, structured records or nothing, and record that decision in the corpus documentation. Under the SEC's 2018 rule, phased in by filer category from 2019, operating companies embed XBRL tags directly in the HTML filing instead of filing a separate XBRL exhibit [7][8]. That means the primary document is an iXBRL HTML file with ix:nonFraction and ix:nonNumeric elements wrapped around values and text blocks, plus hidden header content.

Naive HTML-to-text conversion causes three common failure modes. Hidden ix:header content and context definitions leak into the text; table cells collapse into unordered number strings; and text-block tags around whole notes produce duplicated or truncated passages. The SEC staff's EDGAR XBRL Guide (August 2026 edition, as of October 2026) documents the filing structure your parser needs to respect [9].

Practical options, roughly in order of effort: strip tables entirely and keep prose; serialize tables to Markdown or a pipe format with row and column headers preserved; or keep tagged facts as structured side data (concept name, period, unit, value) linked back to the passage that discusses them. The last option is the most useful for numeric reasoning tasks and the most work to validate. For extracting statements from scanned or non-EDGAR PDFs, see financial statement spreading data.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy you need it
accession_number0000000000-24-000000Stable ID for dedup and leakage checks
cik / form_type0000000000 / 10-KFilter by issuer and form
period_of_report / filed_at2023-12-31 / 2024-02-20Time-based splits and recency weighting
sectionitem_7_mdnaSection-level filtering
textProse with tables removedTraining payload
tables_policydropped / markdown / structuredDocuments extraction choice
xbrl_facts_refpointer to side tableLinks tagged values to prose
minhash_cluster_idc_48211Near-duplicate group
parser_versionedgar-extract 2.3.1Reproducibility

For a broader list of fields to require from any text supplier, see metadata fields for licensed text corpora.

How much deduplication filings need

Filings need both exact and near-duplicate removal, at the passage level as well as the document level. Language-model datasets in general contain many near-duplicates and long repeated substrings, and deduplication reduces verbatim memorization in the trained model [10]. Filings are an extreme case because issuers copy prior-year language forward and law firms reuse templates across clients.

The reported effect can look small at document level: in the 400M-token filings corpus above, MinHash deduplication removed about 1.9% of tokens [3]. That figure reflects a corpus already filtered to narrative sections; run on raw full filings, with exhibits and boilerplate included, expect a larger share. A practical setup is document-level MinHash with locality-sensitive hashing, then paragraph-level exact hashing to catch safe-harbor and certification text; the mechanics are covered in near-duplicate detection with MinHash and LSH.

Also decide how to treat year-over-year near-duplicates within one issuer. Removing them shrinks the corpus but keeping one version per cluster loses the signal of what changed, which is often the most informative text in a 10-K. One compromise is to keep the latest version in full and keep only changed paragraphs from earlier years.

What share of a financial mixture filings should take

Filings should be one component of a domain mixture, not the whole of it. In one published continued-pretraining setup, filings contributed 3.3B words, about 16.5% of the financial mixture, with the rest drawn from financial news [4]. Filings give formal disclosure language; news, analyst commentary and transcripts add the conversational and event-driven register filings lack.

Continued pretraining on a narrow register also risks degrading general capability, so most teams replay some general-domain data alongside the domain corpus. Treat the 16.5% figure as one data point, not a recommendation; tune the ratio against held-out financial tasks and a general benchmark at the same time. For planning the overall domain corpus, see sourcing domain corpora for continued pre-training.

Using filings for long-context training

Filings are long, internally cross-referenced documents, which makes them good long-context material if the reconstruction preserves document order and layout. Recent work argues for layout-faithful reconstruction of long filings so that section headings, notes and table positions survive into the training sequence [5]. Itemized JSON corpora are convenient for section filtering but break exactly the cross-references (for example, MD&A pointing to Note 12) that long-context training needs.

Long-context evaluation should probe the middle of the document, not only the start and end, since models tend to use information at the edges of long contexts more reliably [11]. A useful internal test is to ask about a figure that appears only in a footnote deep in a 10-K and check whether the answer cites the right note. Keep two builds of the same filings: a section-split build for continued pretraining and a whole-document build for long-context work.

Keeping filings-based benchmarks out of training data

Benchmarks built on SEC filings are evaluation assets, and any filings corpus you train on can contaminate them. Fin-RATE, for example, builds real-world financial analytics and tracking tasks on SEC filings [6]. If your pretraining corpus contains the same accession numbers, your benchmark score measures recall, not reasoning.

Run a leakage check before every training run:

  • Collect accession numbers and CIK-period pairs used by each evaluation set.
  • Remove those filings, and their amended versions (10-K/A), from the training corpus.
  • Run n-gram overlap between benchmark passages and the deduplicated corpus to catch copies in other forms or in news coverage.
  • Prefer time-based splits: train on filings before a cutoff date, evaluate on filings after it.

Rights and provenance questions for EDGAR-derived corpora

EDGAR filings are public records, but a packaged corpus still comes with its own terms and provenance gaps. Check the license on the derived release itself (Hugging Face card, repository license, vendor terms), whether exhibits authored by third parties are included, and whether the builder documented the crawl date and parser. Whether a given release permits commercial training is a contract question, not a property of the underlying filings; pre-training data license rights covers the grant language to look for.

Public filings will also overlap with what is already in web crawls. If novelty matters, test candidate data against public pretraining corpora before paying for it, as described in testing licensed data for overlap with public web corpora. The broader text cluster, including news archives and patents, is mapped in the text datasets hub.

When filings are not enough for your financial model

Filings rarely cover the operational side of finance work that agents and copilots are asked to perform. SourceX sources operational datasets from US companies on request, including finance and legal workflows, documents and support or sales histories, and manages the licensing process. These are not held in stock, and a request does not guarantee a match; finance-sector buyers can see how requests are framed on the finance buyers page or describe the data at sourcex.si/buyers.

Sourcing financial operational data beyond SEC filings

If your financial model needs proprietary workflow records that EDGAR cannot supply, describe the data rather than the companies you want it from. SourceX looks for US businesses that hold it, rights-reviews each dataset for ownership and consents, and delivers only under a license that defines records, uses, term and delivery, after the supplier approves the release. Describe your financial data requirement to SourceX.

Sources

  1. arXiv (Loukas et al.), "EDGAR-CORPUS: Billions of Tokens Make The World Go Round" (2021). https://arxiv.org/pdf/2109.14394
  2. Daft (Eventual), "SEC EDGAR case study". https://www.daft.ai/blog/sec-edgar-case-study
  3. arXiv, "The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining" (2025). https://arxiv.org/pdf/2512.12384
  4. arXiv, "Efficient Continual Pre-training for Building Domain Specific Large Language Models" (2023). https://arxiv.org/pdf/2311.08545
  5. papers.cool (arXiv listing), "Layout-faithful long-document reconstruction (arXiv 2606.18192)" (2026). https://papers.cool/arxiv/2606.18192
  6. arXiv, "Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings" (2026). https://arxiv.org/pdf/2602.07294
  7. U.S. Securities and Exchange Commission, "Inline XBRL". https://www.sec.gov/data-research/structured-data/inline-xbrl
  8. U.S. Securities and Exchange Commission, "Inline XBRL Filing of Tagged Data (Release No. 33-10514)" (2018). https://www.sec.gov/files/rules/final/2018/33-10514.pdf
  9. U.S. Securities and Exchange Commission staff, "EDGAR XBRL Guide" (2026). https://www.sec.gov/files/edgar/filer-information/specifications/xbrl-guide.pdf
  10. arXiv (Lee et al., ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  11. arXiv (Liu et al.), "Lost in the Middle: How Language Models Use Long Contexts" (2023). https://arxiv.org/pdf/2307.03172

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data