Skip to content

Data quality, coverage and contamination

Quality Filtering for Pretraining-Scale Text: Heuristics, Classifiers and Their Side Effects

Quick answer

Pretraining data quality filtering removes or down-weights documents using three families of signals: cheap heuristic rules (length, symbol ratios, repetition, language ID), reference-model perplexity, and learned quality classifiers such as FineWeb-Edu's educational-value scorer or QuRating's pairwise-trained raters. Filters improve benchmark scores, but they also shift domain mix, can raise toxic generation, and often penalize valuable business text. Choose filters by ablating them against your own target tasks, not by default thresholds [1][2].

By SourceX Editorial · Updated

What each filter family actually measures

Each filter family encodes a different definition of "quality," and that definition decides which documents survive. Heuristic rules detect malformed or boilerplate text; perplexity filters detect text that looks unlike a reference corpus; learned classifiers detect text that resembles a chosen positive set or satisfies a rubric. None of them measures usefulness for your model directly.

Heuristic rules are document-level tests in the style popularized by C4 and Gopher-era pipelines: word count bounds, mean word length, ratio of alphabetic characters, fraction of lines ending in punctuation, bullet and ellipsis density, duplicated n-gram fractions, stop-word presence and a language-ID confidence score (fastText lid.176 is the common choice). They are cheap, explainable and auditable, which is why nearly every pipeline runs them first. Their failure mode is structural bias: tables, code, log lines, forms and transcripts break "prose" assumptions and get dropped.

Perplexity filtering scores each document with a small language model (often an n-gram KenLM model trained on a reference corpus such as Wikipedia) and keeps the low- or mid-perplexity buckets. It is fast at trillion-token scale, but it rewards similarity to the reference domain rather than correctness. Specialized vocabularies (part numbers, legal citations, ICD codes, ticket shorthand) inflate perplexity even when the text is precise.

Learned quality classifiers are trained on a positive set and a negative set, or on LLM-generated ratings, and then applied to every document. Longpre et al. note that the LLaMA, GPT and PaLM model families relied on quality classifiers of this kind [1]. DCLM made model-based filtering a central variable in its controlled comparison of curation strategies at fixed compute [2], and newer benchmarks test LLMs directly as rating and preparation tools [3].

How QuRating and FineWeb-Edu turn LLM judgments into filters

Both methods distill expensive LLM judgments into a cheap scorer that can run over a full corpus. They differ in how the judgments are elicited and in how the scores are used to select data.

QuRating asks an LLM to compare pairs of documents on named criteria (writing style, required expertise, facts and trivia, and educational value) and then trains a smaller rater to output a scalar score for every document; recent benchmarking work treats it as a reference example of per-example LLM scoring for data selection [3]. Pairwise comparison is used because absolute scores from a single prompt tend to be noisy. A practical lesson from this line of work is that keeping only the top-rated documents can collapse diversity, so many teams sample with scores as weights instead of applying a hard cut.

FineWeb-Edu applies the same distillation idea with an educational-value rubric: a large instruction-tuned model rates a sample of web pages on a 0-5 scale, and a small embedding-based regressor then scores the full corpus. The rubric deliberately favors explanatory, school-level content, which is a design choice, not a neutral measure of value. Business documents rarely read like lessons, so expect them to cluster at the low end of such scales.

The general pattern, described in data-centric surveys, is to combine metric checks (format, duplication, diversity, fluency, factual accuracy) with LLM-based scoring and human review of samples [4]. For rater reliability and bias testing, see using an LLM judge to score training data.

Side effects: what filtering does to the corpus and the model

Quality filters change the distribution of the corpus, so they change the model in ways benchmark gains do not reveal. The controlled evidence comes from Longpre et al., who pretrained 1.5B-parameter models on variants of the same corpus [1].

  • Capability and toxicity can rise together. Quality filtering removed more than 10% of the data and improved downstream performance, yet it also increased toxic generation [1]. A quality filter is not a safety filter; run toxicity filtering and its trade-offs as a separate, measured decision.
  • Effects are not predictable from domain labels. The same filter helped some tasks and domains and hurt others, and its effect could not be inferred from the domain's characteristics [1]. Tune per target task.
  • Diversity loss. Hard top-k thresholds concentrate the corpus on one style; weighted sampling preserves more diversity.
  • Rubric bias. An educational-value rubric favors explanatory prose. Procedures, contracts, support tickets, changelogs and spreadsheet-derived text describe how work is done but rarely read like a textbook.
  • Dialect and register skew. Classifiers trained on encyclopedic or curated positives can drop informal, regional or non-native English at higher rates, which narrows coverage of the deployment population.

Our working hypothesis for business corpora: operational text from support, engineering, finance and legal workflows will often score low on web-tuned classifiers while being exactly the signal a domain model needs. Treat this as something to test on each corpus, not as an established result.

Where to place filters in a licensed-corpus pipeline

Filters belong after extraction and normalization and alongside deduplication, with every stage logged so the removal can be explained. The order changes results: deduplicating after quality scoring wastes compute on documents you will drop, while scoring before boilerplate removal penalizes good documents for their headers and footers.

A practical order for licensed text:

  1. Format extraction and normalization (PDF, HTML, DOCX, email MIME to UTF-8 text; strip signatures and templated disclaimers).
  2. Language ID and encoding checks.
  3. Exact and near-duplicate removal; see MinHash and LSH deduplication.
  4. Heuristic rules tuned per source type, not one global set.
  5. Model-based quality scoring, stored as a score rather than applied as a hard cut.
  6. Mixture weighting and sampling at training time, informed by ablations.

Storing scores instead of deleting documents keeps the decision reversible. It also lets you report, per source, how many tokens each stage removed, which belongs in a dataset quality report. Corpus-scale filtering is a different problem from curating a few thousand SFT pairs; for that, see instruction-tuning data quality filtering.

Validating a filter on licensed business text

A filter is validated when ablations on your target tasks beat an unfiltered or uniformly sampled baseline at equal token budget. Classifier accuracy against its own labels says nothing about downstream value.

Run small-scale proxy ablations (for example, a 1B-class model on a fixed token budget) for each filter variant, and evaluate on held-out tasks that match the intended use: domain QA, document extraction, ticket classification, code or log understanding. Inspect the removed set by stratified sample, not just the kept set, and break removal rates out by source system, document type and date. If a classifier removes most of one source, decide deliberately whether that source is noise or the reason you licensed the corpus.

Illustrative example: invented to show structure; it does not describe an available dataset.

Filter stageSignal and threshold (example)Removed share (example)What to inspectTypical false positive in business text
Language IDfastText lid.176 English score < 0.652.1%Mixed-language ticketsPart-number-heavy lines misread as other languages
Heuristic: alpha ratioAlphabetic chars < 70% of tokens6.4%Tables, CSV exportsInvoice line items and BOM tables
Heuristic: repetitionDuplicate 3-gram fraction > 0.203.8%Templated emailsRunbooks with repeated step headers
PerplexityKenLM (Wikipedia) top-decile perplexity9.7%Domain jargonLegal citations, ICD-10 codes, error logs
Quality classifierEdu-style score < 2 of 531.5%Removed-set sample by sourceResolved support threads, change records
DecisionWeighted sampling by score; floor per sourcen/aPer-task ablation deltasn/a

Request these numbers from any supplier that pre-filtered the corpus, along with the classifier version, prompt or rubric, and the unfiltered remainder if the license allows. If the supplier already applied a web-tuned classifier, you may be paying for the documents least likely to help your domain. Check too whether the licensed text already appears in public crawls; see testing novelty against pretraining corpora.

Questions to settle before licensing a large text corpus

Settle filtering questions in the commercial terms, because the answers decide what you receive. Ask:

  • Has the supplier filtered, deduplicated or rewritten the text, and with which tools, versions and thresholds?
  • Can you receive raw text plus per-document metadata (source system, document type, timestamps) so you can run your own filters?
  • Are personal details removed or replaced, and does that process alter perplexity or classifier scores (placeholder tokens such as [NAME] can shift both)?
  • Do the license's allowed uses cover pretraining and continued pretraining, and do they cover derived scores and filtered subsets? See the rights grant for pre-training.

For sizing and sourcing the corpus itself, see licensed text corpora for LLM pre-training and domain corpora for continued pre-training. The pretraining glossary entry and the data quality hub cover the surrounding concepts. If you need operational text that web-tuned filters underrepresent, you can describe the corpus to SourceX.

Sourcing licensed text corpora for pretraining filters to work on

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages the licensing process. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, with personal details removed or replaced before delivery; nothing is held in stock, and a request does not guarantee a match. Describe the text you need at sourcex.si/buyers.

Sources

  1. Longpre et al., arXiv (NAACL 2024), "A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity" (2023). https://arxiv.org/pdf/2305.13169
  2. Li et al., arXiv, "DataComp-LM: In search of the next generation of training sets for language models" (2024). https://arxiv.org/pdf/2406.11794
  3. arXiv, "DataPrep-Bench: Benchmarking LLMs as Training Data Preparators" (2026). https://arxiv.org/pdf/2607.20465
  4. arXiv, "Towards Next-Generation LLM Training: From the Data-Centric Perspective" (2026). https://arxiv.org/pdf/2603.14712

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data