Skip to content

Data quality, coverage and contamination

Is Licensed Data Already in Public Web Crawls? Testing Novelty Against Pretraining Corpora

Quick answer

To check whether data sold as proprietary is already in Common Crawl or open pretraining corpora, run three tests on a supplier sample before you sign: URL and domain lookups against the Common Crawl index, exact and near-duplicate text search against indexed open corpora such as RedPajama, Dolma or C4, and, as a weak supplementary signal, model-side membership tests. String overlap gives hard evidence; likelihood tests give probabilities. Record the method, corpus versions and overlap rate in the purchase file and price accordingly.

By SourceX Editorial · Updated

Why "proprietary" claims need a novelty test

A novelty test answers one commercial question: how much of what you are paying for has your model, or any model trained on open crawls, already seen. Business data often has a public twin. Help-center articles get published on the vendor's knowledge base, contract templates sit on law-firm blogs, product manuals are posted as PDFs, and SEC filings are on EDGAR, so a "support corpus" or "finance document set" can overlap the open web by a meaningful margin even when the supplier honestly holds it as internal records.

The test is distinct from two neighboring checks. Overlap with data you already hold is covered in checking overlap between a new dataset and data you already own, and removal of benchmark items is covered in decontaminating a licensed training set against public benchmarks. This page is only about the open web and open corpora. For background on why non-public operational records add something crawls lack, see proprietary text data beyond web crawls.

Common Crawl itself states that it follows robots.txt and avoids paywalled and login-protected sites [4]. That is the core reason genuinely internal records (ticket threads behind an agent console, CRM notes, private engineering wikis) should show near-zero overlap. When a sample shows high overlap, either the content was published somewhere, or the "proprietary" set was assembled from public pages.

Test 1: URL and domain lookup against the Common Crawl index

The fastest test is a URL lookup, and it only works when records carry source URLs or the supplier names the domains its content was published on. Common Crawl exposes a CDX index server (index.commoncrawl.org) per crawl, keyed by URL, and a columnar index in Parquet that can be queried with Athena, Spark or DuckDB by url_host_registered_domain, url_path, fetch_status and content_mime_type. A hit tells you the page was captured in that crawl, not that it survived the filtering of any specific pretraining set.

Run the lookup across several crawl IDs, because each crawl captures a different sample of pages. Then repeat the query for downstream filtered corpora where possible, since a page in raw WARC files may be dropped by quality filters described in pretraining text quality filtering. If the supplier's records have no URLs, as most internal exports do, skip to text search.

Common failure mode: the supplier's public help center lives on a third-party domain (a hosted knowledge-base or docs platform subdomain), so a lookup on the company's own domain returns nothing while the content is fully crawled. Ask for every domain the content was ever published on.

Test 2: Exact and near-duplicate text search against open corpora

Text search is the core test because it measures overlap directly from content, with no URLs needed. Two families of methods apply, both described for corpus-scale deduplication by Lee et al. [2]: exact substring matching (suffix arrays find any shared span above a token threshold, such as 50 tokens) and near-duplicate matching (MinHash signatures over word or character shingles, bucketed with locality-sensitive hashing, flagged above a Jaccard threshold such as 0.8). The same paper shows how repetitive open corpora are, with one sentence in C4 repeated over 60,000 times [2], so boilerplate will match even when the substance is new.

You do not need to rebuild a trillion-token index. Research tools such as infini-gram publish suffix-array indexes over open corpora such as RedPajama and Dolma that return n-gram counts quickly, the same exact-substring approach Lee et al. used for deduplication [2], which makes span-level queries of a sample practical. For MinHash, you can stream a public corpus shard by shard and probe your sample's signatures; details are on near-duplicate detection with MinHash and LSH.

Exact and near-duplicate matching both miss paraphrase. LMSYS showed that rephrased test items evade simple n-gram checks, and used an embedding search plus an LLM judge to find them when running overlap checks against public corpora such as The Stack and RedPajama [1]. For a high-value purchase, add an embedding nearest-neighbor pass over a sampled public corpus and have a model or reviewer judge the top matches.

A cheap complement is manual web search: pick 30 to 50 distinctive sentences (rare product names, specific error strings, unusual phrasing) and search them in quotes. It catches content that was published after the open corpora you indexed were frozen.

Test 3: Model-side membership signals, and their limits

Membership tests ask a model whether it has likely seen a text, and they return probabilities rather than proof. Membership-inference and log-probability signals, such as Min-K% Prob, which averages the log-likelihood of a document's lowest-probability tokens, are used as complements to string matching rather than replacements [3]. Their published accuracy is well short of certainty, and likelihood scores are easily confused by text full of common, high-probability words, which is exactly what templated support replies and boilerplate contracts look like.

Treat these signals as supplementary, as market practice also frames them [3]. Use them on a small held-out set against an open-weights model whose training corpus you can partially check, and compare against a control set you know is unseen, such as records dated after the model's cutoff. Never reject or accept a dataset on a membership score alone.

For closed models, the public training-content summaries that general-purpose model providers must publish under Article 53(1)(d) of the EU AI Act, using the Commission template dated 24 July 2025 [5], describe data sources at a high level. They can tell you a provider used public crawls; they will not tell you whether a specific document was included.

Test plan to run on a supplier sample

A useful novelty test fits in a short protocol you agree with the supplier before the sample ships. Use the table below as a starting template; thresholds are examples to tune, not standards.

Illustrative example: invented to show structure; it does not describe an available dataset.

StepInputMethodOutput metricExample flag
1. ScopeSupplier's list of publication domainsCommon Crawl CDX or columnar index, 3+ crawl IDs% of records with a URL hit> 5% of records
2. Span overlap2,000-record random sampleInfini-gram or suffix-array query, 50-token spans% of records with any shared span> 10%
3. Near-duplicatesSame sample, boilerplate strippedMinHash, 5-gram shingles, Jaccard 0.8% of records with a near-duplicate> 5%
4. ParaphraseTop 200 records by step 2–3 scoreEmbedding search plus LLM or human judge% judged rephrased public contentAny cluster from one public source
5. Manual search40 distinctive sentencesQuoted web searchCount of exact public hits> 3 hits
6. Membership (optional)500 records plus post-cutoff controlMin-K% Prob on an open-weights modelScore gap vs. controlReport only

Strip boilerplate before step 3: signatures, legal footers, greeting templates and auto-generated ticket headers will match the web by design. Then read every high-overlap cluster by hand to name its likely public origin. For sample sizing, how many records to check gives the arithmetic for estimating an overlap rate with a confidence interval.

Expected overlap by data type, and what it means for price and scope

Some public overlap is normal in business data, so the goal is to measure it and scope around it, not to demand zero. Partial overlap is expected wherever a company publishes part of its operations, as this hypothesis table summarizes.

Illustrative example: invented to show structure; it does not describe an available dataset.

Data typeLikely public twinExpected overlapBuyer response
Support ticket threadsPublished help-center articles quoted in agent repliesLow to moderateKeep threads; exclude or discount KB article text
Sales call notes, CRM recordsRarely publicVery lowHigh overlap is a red flag; ask how the set was built
Engineering tickets and code reviewPublic issue trackers for open-source componentsLow to moderateSeparate public-repo items; check licenses on those
Contracts and legal templatesPublic form contracts, filed exhibitsModerateValue lies in negotiated versions and redlines
Finance documentsSEC filings, earnings transcriptsModerate to high for reporting; low for internal workflowScope to internal workflow records

If overlap is concentrated in an identifiable public slice, carve it out of the license scope or the price rather than walking away. Overlap also raises a rights question: third-party text inside a licensed corpus may carry its own terms, discussed in third-party content inside licensed corpora.

What to record in the purchase file

Record the novelty test so the result can be reproduced and defended later in model documentation or an audit. A short structured entry is enough.

Illustrative example: invented to show structure; it does not describe an available dataset.

novelty_test:
  dataset_ref: "support-threads-sample-v2"
  sample: { records: 2000, selection: "uniform random by ticket_id", seed: 1729 }
  corpora_checked:
    - { name: "Common Crawl", crawl_ids: ["<three crawl IDs>"], method: "columnar index, registered domain + path" }
    - { name: "RedPajama", method: "infini-gram exact span, min 50 tokens" }
    - { name: "C4", method: "MinHash 5-gram, 128 permutations, Jaccard >= 0.8" }
  results:
    url_hit_rate: 0.012
    span_overlap_rate: 0.064
    near_dup_rate: 0.021
    dominant_public_source: "vendor help-center articles quoted in replies"
  membership_signal: { method: "Min-K% Prob, k=20", status: "reported only, not decisive" }
  decision: "accept with KB-article text excluded from scope"
  reviewer: "<name>"
  date: 2026-10-09

Store the corpus versions and thresholds with the numbers; an overlap rate without its method cannot be compared across suppliers. Add the entry to the evidence pack described in the AI training data due diligence checklist, and see licensed vs synthetic vs scraped data for why provenance class matters as much as content.

How SourceX handles novelty for sourced data

SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records, documents, and finance and legal workflows, and it does not source scraped web content. Because records come from inside supplying businesses, the categories most likely to be novel are the ones crawlers cannot reach, but you should still run your own novelty test on any sample. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and diligence materials on source and preparation are prepared per dataset, which gives your novelty test a documented starting point. Teams can describe the data they need to SourceX; a request does not guarantee a match. More quality topics are collected on the data quality cluster hub and the AI data hub.

Sourcing non-public data for your model

If your test plan shows that what you need is not in public crawls, SourceX can look for US businesses that hold the data you describe, with every release approved by the supplying company. Nothing is contracted until a supplier agrees, and pricing and allowed uses are set in a license per deal. Describe the non-public data you need.

Sources

  1. LMSYS Org, "Catch me if you can! How to beat GPT-4 with a 13B model (LLM Decontaminator)" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
  2. Lee et al., ACL 2022 (arXiv:2107.06499), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  3. Maxim AI, "LLM data quality". https://getmaxim.ai/blog/llm-data-quality
  4. Common Crawl Foundation, "Common Crawl's Submission to the UK's Copyright and AI Consultation". https://commoncrawl.org/uk-copyright-and-ai-consultation
  5. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data