Skip to content

Fine-tuning and post-training data

Sourcing domain corpora for continued pre-training

Quick answer

A continued pretraining dataset is a large body of unlabeled, in-domain text, such as filings, contracts, service manuals, engineering records or support histories, used to keep training an existing base model on its next-token objective before any fine-tuning. To source one, find the systems that hold the text, count usable tokens with your own tokenizer after deduplication, test how much of it the base model has already seen, and license it with rights that cover training and keeping the resulting weights.

By SourceX Editorial · Updated

When a domain corpus is the right purchase

Buy a domain corpus when the base model lacks a field's knowledge and language, not when it only needs a task or an output format. The LIMA authors fine-tuned a 65B-parameter model on 1,000 curated examples and concluded that almost all knowledge comes from pre-training, while a small instruction set mainly teaches format [1]. Continued pre-training on domain data, often called domain-adaptive pre-training (DAPT), adds knowledge: the model reads raw documents with loss on every token, with no prompts, answers or preference labels.

Three signals favor it over fine-tuning or retrieval:

  • High perplexity on held-out domain documents. Score a supplier's sample with your base model; text it already predicts well has little left to teach.
  • Vocabulary the model fragments or misuses. Part numbers, billing codes, statute citations and internal acronyms that split into many tokens or come back wrong.
  • Stable knowledge absent from public text. Facts that change weekly or must be cited suit retrieval better; see what data fine-tuning, RAG and continued pre-training each need.

High-quality domain text added late in a pre-training run is a different purchase, covered in mid-training and annealing data. Labeled records for supervised training are covered in training data for domain-specific fine-tuning.

Where in-domain text originates, by sector

Domain text comes from public records, which are easy to obtain but probably already in your base model's training data, and from operational records inside companies, which carry new information along with heavier rights and privacy work.

SectorPublic or openly licensed textRecords held inside companiesWhat to check
FinanceSEC filings; EDGAR-CORPUS holds 10-K reports from 1993 to 2020, split into items as JSON [2]Credit memos, research notes, loan files, reconciliation commentaryTables and tagging markup; one study built a 400-million-token corpus from 10-K, 10-Q and DEF 14A filings, keeping narrative sections such as MD&A and Risk Factors [3]
LegalOpinions, filings, statutes, regulations; Pile of Law assembles about 256 GB from 35 sources [4]Contracts, briefs, matter memos, redlinesPrivilege; editorial layers on public text, such as headnotes, can carry their own copyright [5]
EngineeringProduct documentation whose terms allow trainingService manuals, change orders, design reviews, postmortems, ticketsVendor manuals in a company archive may belong to the vendor; one chip-design project combined tool manuals, engineer Q&A, papers and script documentation, over 200,000 pages [6]
Customer operationsLimitedSupport tickets, chat transcripts, call notesDense personal data; macros create near-duplicates
HealthcareLimited; most clinical text is protected health informationClinical notes, prior-authorization lettersHIPAA de-identification by Safe Harbor (18 listed identifiers) or Expert Determination [7]

Filings and court opinions are widely crawled, so check novelty before paying to clean them, using tests of licensed data against public web crawls. Openly licensed collections still serve as a baseline and as replay text: the Common Pile v0.1 gathers 8 TB of openly licensed and public-domain text from 30 sources, and its authors report 7B models trained on it that compete with models trained on unlicensed text at similar compute [8]. Verify each license anyway; an audit of more than 1,800 text datasets found license omissions above 70% and error rates above 50% on popular hosting sites [9].

Operational records are likely to hold most of the new information. SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, on request rather than from stock, so a request does not guarantee a matching dataset; it does not source scraped web content. Buyers describe the domain text they need to SourceX without approaching companies themselves, and every release is approved by the supplying company. Document-heavy domains are covered in enterprise document datasets for AI training.

Qualify the corpus with a token audit before you sign

The number that matters is usable tokens: tokens counted with your own tokenizer after extraction, boilerplate removal and deduplication, broken out by source system, document type and year. Gigabytes, pages and word counts overstate trainable volume by an unknown margin until measured. Request the audit below and reproduce it on a sample; budget sizing is covered in how much data fine-tuning and continued pre-training need.

MeasureWhy it mattersWhat to request
Usable tokens under your tokenizerFertility (tokens per word) varies by tokenizer and language; in one adaptation study, replacing about 5,000 tokens, 10% of the vocabulary, improved fertility by 42% for Hungarian and 73% for Thai [10]Count by named tokenizer, per source and year
Exact and near-duplicate rateLee et al. found one sentence repeated over 60,000 times in C4, and deduplicated training emitted memorized text about ten times less often [11]; the filings corpus above lost about 1.9% of tokens to MinHash deduplication [3]Method (exact hash, MinHash with LSH), threshold and removal rate per source
Boilerplate shareSignatures, disclaimers, quoted reply chains and page headers repeat across thousands of documentsCharacters removed per document and the rules
Language mixBusiness archives can mix languages; fastText-based identifiers such as GlotLID give confidence scores usable as filter thresholds [12]Share by language, tool and threshold
Time and type mixTerminology changes, recent years may be thin, and auto-generated notices can dominate an exportTokens by year and by document type
Extraction qualityScans and tables produce OCR noise and broken reading orderMethod, OCR confidence and the hardest documents
HeadroomText the base model already predicts well adds littleA random sample for perplexity and overlap tests

Ask for deduplication per source, because a low overall rate can hide a ticket export made mostly of macros. Mechanics are in near-duplicate detection with MinHash and LSH and quality filtering for pretraining-scale text; terminology coverage is in sourcing text rich in industry jargon.

What each delivered document should carry

Each document should arrive as one cleaned text record with a stable ID and enough metadata to filter, trace and remove it. IDs let you drop a source after a rights dispute, hold out a time-based evaluation slice and answer disclosure questions about training data.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "doc_id": "eng-eco-000418273",
  "source_system": "plm_change_orders",
  "doc_type": "engineering_change_order",
  "created_at": "2021-03-14",
  "language": "en",
  "lang_confidence": 0.98,
  "extraction": "native_pdf_text",
  "tokenizer": "buyer_named_tokenizer",
  "token_count": 4310,
  "dedup": {"method": "minhash_lsh", "threshold": 0.8, "cluster_id": "c-55120", "kept": true},
  "boilerplate_chars_removed": 1260,
  "pii": {"method": "ner_regex_surrogates", "entities_replaced": 12},
  "license_id": "LIC-0042",
  "third_party_content": false,
  "split": "train",
  "text": "ECO 2021-114: Replace bracket P/N 7730-22 with 7730-24 on line 3. Reason: field reports of fatigue cracking at the weld toe. Approved by [PERSON_0193], Manufacturing Engineering."
}

The record is pretty-printed here; in a JSON Lines file each record is one UTF-8 JSON value on its own line, and gzip compression (.jsonl.gz) is recommended [13]. Keep every field except text in a separate manifest keyed by doc_id, so filtering never requires reading the corpus. For dataset-level documentation, Croissant describes datasets, files and record structure in JSON-LD built on schema.org [14]. Shards and transfer are covered in dataset delivery formats and transfer.

Scope privacy before redaction starts

Free text puts identifiers anywhere in a document, and models trained on it can reproduce them. Carlini et al. extracted hundreds of verbatim training sequences from GPT-2, including contact details, and found larger models more vulnerable [15]. Detection tools reduce that risk without removing it; the Presidio project cautions that, because it relies on trained ML models, there is no guarantee it finds all sensitive information [16].

Decide at corpus scale, before any redaction pass runs:

  • Exclude by source, not by sentence. Leave out HR, payroll and medical correspondence at the system level when the domain does not need them.
  • Choose surrogates deliberately. A bare [NAME] repeated across a corpus loses who-did-what and can become a pattern the model reproduces; indexed tags such as [PERSON_0193] keep coreference, while realistic substitute names avoid the pattern but can be mistaken for real people.
  • Measure residuals per source. Ask for the share of sampled documents still containing an identifier.
  • Treat health records separately. HIPAA Safe Harbor removes 18 listed identifiers of the individual and of relatives, employers and household members; Expert Determination relies on an expert finding that identification risk is very small [7].

SourceX removes or replaces personal details such as names, emails, phone numbers and account numbers before delivery, records the method for each dataset and checks a sample afterward; no de-identification method is perfect. For health records it requires HIPAA de-identification by Safe Harbor or Expert Determination before anything is considered for a license. Measurement methods are in PII redaction for LLM training data.

License terms a continued pre-training corpus needs

A license for a continued pre-training corpus must grant training on the full text, not only fine-tuning on derived examples, let you keep and run the resulting weights after the term, and cover the artifacts your pipeline creates. Fine-tuning grants often stop short; compare fine-tuning-only data licenses with the rights grant pre-training needs.

  1. Use grant. Does it name continued pre-training, and does it reach models you later fine-tune or distill from the adapted checkpoint?
  2. Weights after term. Do deletion or termination clauses reach checkpoints and deployed weights, or only corpus copies?
  3. Derived artifacts. Tokenized shards, deduplication indexes, embeddings and synthetic text generated from the corpus.
  4. Distribution. Hosted API only, or open weights; memorization risk makes this a real difference for proprietary text.
  5. Third-party content. Vendor manuals, attached articles and customer data inside records may not be the supplier's to license.

The fifth point matters most for corpora built on public text. On 29 September 2026 the US Court of Appeals for the Third Circuit held in Thomson Reuters v. ROSS Intelligence that the Westlaw headnotes at issue were copyrightable and that using them to train ROSS's non-generative legal-research tool was not fair use [5]. The US Copyright Office's Part 3 report on generative AI training, still a pre-publication version as of October 2026, says copying works into training datasets may be prima facie infringing unless an exception such as fair use applies [17].

Disclosure duties shape what to request. California AB 2013 requires developers of generative AI systems available to Californians to post training-data documentation covering items such as dataset sources or owners, copyrighted or licensed material, personal information, collection periods and synthetic data; it was due by 1 January 2026 and is required again for substantial modifications [18]. If the adapted model is placed on the EU market as a general-purpose AI model, its provider must keep a copyright policy and publish a training-content summary on the AI Office template under Article 53 of the AI Act [19]. Ask counsel whether your run is a substantial modification under AB 2013 or makes you the adapted model's provider under the AI Act.

Every SourceX dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place, and is delivered under a license defining the records included, permitted uses, term and delivery. Diligence materials on source, rights, preparation and allowed use are prepared per dataset for buyer review; clause definitions are in AI data license terms explained.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Plan the replay mix with the purchase

Domain text alone erodes general abilities, so budget general "replay" text in the same purchasing cycle and check its license too. A study adapting an English model to Hungarian and Thai reports that naive adaptation causes catastrophic forgetting and uses a data-mixing recipe to keep English regressions small [10]. The domain side can itself be a mix: a financial continual pre-training study combined 3.3 billion words of SEC filings, 16.5% of its corpus, with financial news drawn from Common Crawl [20].

Red flags in a domain corpus offer

These signs mean a corpus will cost more to prepare, or carry more risk, than the offer suggests:

  • Mostly public filings or court opinions, priced as proprietary.
  • A rights basis of "publicly available" or "fair use" instead of a grant from the rights holder.
  • No document IDs, so a source cannot be traced or removed.
  • "PII removed" with no method, residual measurement or sample.
  • Weights rights that end with the license term.
  • Synthetic or machine-translated text without a flag, although AB 2013 documentation asks whether synthetic data was used [18].

Checklist for a continued pre-training corpus request

Send these items so every supplier quotes against the same specification:

  • Domain, document types and source systems, plus exclusions
  • Date range and minimum tokens for recent years
  • Languages and language-ID threshold
  • Minimum usable tokens under your named tokenizer after cleaning
  • Deduplication and boilerplate rules, reported per source
  • Extraction method for PDFs and scans, with OCR confidence
  • De-identification scope, method, surrogate style and residual sample
  • Compressed JSON Lines shards plus a manifest with the fields above
  • A sample sized for perplexity and crawl-overlap tests
  • Rights for continued pre-training, derived models, weights after term, artifacts and distribution
  • Documentation answering AB 2013 and AI Act Article 53 training-data questions

Describe the domain corpus you need

At the SourceX buyer page you can submit a data request describing the domain text you need, using the document types, date range, token targets and licensed uses from the checklist above. SourceX looks for US companies that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Request a domain corpus through SourceX.

Sources

  1. Zhou et al., Meta AI and collaborators, "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  2. Loukas et al., "EDGAR-CORPUS: Billions of Tokens Make The World Go Round" (2021). https://arxiv.org/pdf/2109.14394
  3. arXiv preprint 2512.12384, "The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining" (2025). https://arxiv.org/pdf/2512.12384
  4. Henderson et al., "Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset" (2022). https://arxiv.org/pdf/2207.00220
  5. U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (precedential opinion)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
  6. arXiv preprint 2604.27415, "ChipLingo: A Systematic Training Framework for Large Language Models in EDA" (2026). https://arxiv.org/pdf/2604.27415
  7. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  8. Kandpal et al., "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
  9. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version in Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
  10. Csaki et al., "Efficiently Adapting Pretrained Language Models To New Languages" (2023). https://arxiv.org/pdf/2311.05741
  11. Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021; ACL 2022). https://arxiv.org/abs/2107.06499v1
  12. Kargaran et al., "GlotLID: Language Identification for Low-Resource Languages" (2023). https://arxiv.org/pdf/2310.16248
  13. jsonlines.org, "JSON Lines". https://jsonlines.org/
  14. Akhtar et al., MLCommons Croissant working group, "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  15. Carlini et al., "Extracting Training Data from Large Language Models" (USENIX Security 2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
  16. Microsoft (microsoft/presidio project), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  17. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  18. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  19. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  20. arXiv preprint 2311.08545, "Efficient Continual Pre-training for Building Domain Specific Large Language Models" (2023). https://arxiv.org/pdf/2311.08545

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data