Skip to content

Data quality, coverage and contamination

Exact and Near-Duplicate Detection in Licensed Text Datasets with MinHash and LSH

Quick answer

Run deduplication in two passes. First remove exact duplicates with a content hash over normalized text, then find near-duplicates by shingling each document, computing MinHash signatures (128 permutations is a common start), and using LSH banding to surface candidate pairs above a Jaccard threshold near 0.8. Verify candidates, cluster them, and report exact and near-duplicate rates separately. Tune the threshold on hand-labeled pairs from your own corpus, not on defaults.

By SourceX Editorial · Updated

Why duplicate rate belongs in dataset acceptance

Duplicates inflate apparent dataset size, skew the training distribution and raise memorization risk, so a buyer should measure them before accepting a delivery. Lee et al. found that standard language-modeling corpora contained heavy repetition, including a single 61-word sentence repeated more than 60,000 times in C4, and that deduplicating cut the rate at which models emitted memorized training text by roughly a factor of ten [1]. They also showed that duplicates crossing train and validation splits make held-out metrics look better than they are [1].

Licensed operational text has its own duplication patterns, separate from web crawl. Support tickets repeat macro replies and email signatures, email exports carry quoted reply chains, document stores hold the same contract in v3, v3-final and v3-final-signed, and a CRM export may include the same record pulled through two joins. For a priced dataset, a high duplicate rate is also a commercial issue: you pay for records that add little signal. Treat dedup as one input to training data quality metrics and to your broader quality assessment for licensed data.

Exact duplicates: hash normalized text first

Exact deduplication is cheap and should always run before any approximate method. Normalize each record (Unicode NFKC, lowercase if case is not meaningful, collapse whitespace, strip known boilerplate such as signature blocks and ticket-system footers), then compute a SHA-256 or xxHash64 digest and group by digest. Keep the record ID, source system and timestamp on every row so you can trace which copy you kept.

Decide what the unit is before you hash. For tickets, hashing the whole thread misses that two threads share an identical agent reply; hashing each message catches it. For documents, hash both the full document and each paragraph, because paragraph-level exact matches expose templated contracts and copied policy sections that whole-document hashes miss.

Exact substring deduplication goes further. Lee et al. built a suffix array over the corpus and removed repeated spans above a token-length threshold (their ExactSubstr method), which catches long shared passages inside otherwise different documents [1]. It is heavier to run than hashing but is the right tool when the concern is verbatim repetition that a model could memorize, such as legal boilerplate or disclaimers repeated across thousands of records.

How MinHash and LSH find near-duplicates

MinHash estimates Jaccard similarity between documents from short fixed-length signatures, and LSH avoids comparing every pair by bucketing signatures so only likely matches meet [2]. The pipeline has three stages [2]:

  • Shingling. Convert each document to a set of overlapping n-grams. Word 5-grams are a common choice for prose; character 5- to 9-grams work better for short or noisy text such as chat messages or OCR output.
  • MinHash signatures. Apply k hash functions (permutations) to the shingle set and keep the minimum value for each. The fraction of positions where two signatures agree is an unbiased estimate of their Jaccard similarity, and the estimate tightens as k grows.
  • LSH banding. Split the k-value signature into b bands of r rows (b × r = k). Hash each band to a bucket; two documents become a candidate pair if any band matches exactly.

With banding, the probability that a pair with true Jaccard similarity s becomes a candidate is 1 − (1 − s^r)^b. The curve is an S-shape whose steep point sits near (1/b)^(1/r). Choosing b and r moves that point and trades false negatives against candidate volume.

Illustrative example: invented to show structure; it does not describe an available dataset.

Signature setup (k = 128)Approx. threshold (1/b)^(1/r)P(candidate) at s = 0.5P(candidate) at s = 0.7P(candidate) at s = 0.8
b = 16, r = 80.71~0.06~0.61~0.95
b = 32, r = 40.42~0.87>0.99>0.99
b = 8, r = 160.88<0.01~0.03~0.20

The table shows why the band layout matters as much as the nominal threshold. Libraries such as datasketch choose b and r for you from a threshold and weights on false positives and negatives; check what they picked, and always verify candidates by computing the actual Jaccard (or signature agreement) before you merge them.

Starting parameters and how to tune them

A reasonable starting point is 128 permutations and a Jaccard threshold of about 0.8, then tuning on labeled pairs from your own data [3]. Defaults travel badly between corpora: a threshold that works for web pages will merge distinct support tickets that share a long macro, and it will miss contract versions whose only difference is a changed party name and date.

Tune with a small labeled set:

  1. Sample about 300 to 500 candidate pairs across similarity bands (0.5 to 0.6, 0.6 to 0.7, and so on up to 1.0), plus random non-candidate pairs.
  2. Have two reviewers label each pair as duplicate, near-duplicate worth collapsing, or distinct, using a written definition tied to your use (pre-training tolerates more collapse than evaluation sets do).
  3. Plot precision and recall against threshold and shingle size. Pick the threshold where false merges of genuinely distinct records stay acceptable for your task.
  4. Re-run on a holdout sample before applying to the full corpus.

Record the final normalization rules, shingle type and size, k, b, r, threshold and library version in a dedup config file that ships with the cleaned dataset. That config is what lets someone reproduce the duplicate rate months later.

Running dedup at scale with Spark and other engines

MinHash LSH scales well in theory, but common distributed implementations have practical limits you should plan around. A 2026 PyData London talk reported Spark MLlib's MinHashLSH running out of memory on large corpora [4]; a likely cause is the approximate similarity join exploding on oversized buckets.

Failure modes to watch:

  • Hot buckets. Boilerplate-heavy records (auto-replies, empty templates, "Thank you for contacting support") land in the same band buckets and produce quadratic candidate counts. Remove exact duplicates and known boilerplate first, and cap or sample oversized buckets.
  • Shuffle size. Emitting one row per band per document multiplies data volume by b. Store band hashes as 64-bit integers, not strings, and partition by band.
  • Transitive clustering. Candidate pairs form a graph; use connected components (for example GraphFrames or a union-find pass) to build clusters, and inspect the largest clusters, because chained ABC links can merge records that are not similar end to end.
  • Non-determinism. Fix hash seeds and record them, or reruns will produce slightly different clusters and a different duplicate rate.

Some vector databases also expose MinHash LSH as an index type for near-duplicate search, which can suit incremental dedup of ongoing deliveries [5]. For code corpora, where forks and vendored libraries dominate, see deduplicating code training data.

Reporting a duplicate rate a reviewer can trust

A useful duplicate report separates exact from near-duplicates, states the unit and method, and shows the cluster-size distribution instead of a single percentage. A headline "12% duplicates" is uninterpretable without knowing whether it counts records removed, records in clusters, or pairs.

Illustrative example: invented to show structure; it does not describe an available dataset.

dedup_report:
  dataset: support_tickets_delivery_02
  unit: message_body
  records_in: 1_000_000
  normalization: [nfkc, lowercase, collapse_whitespace, strip_signature_blocks]
  exact:
    method: sha256 on normalized text
    clusters: 41_200
    records_removed: 88_500          # keep 1 per cluster
    exact_dup_rate: 0.0885
  near:
    method: minhash_lsh
    shingles: word_5gram
    num_perm: 128
    bands: 16
    rows: 8
    verify: jaccard >= 0.80
    clusters: 23_900
    records_removed: 61_300
    near_dup_rate: 0.0673            # measured after exact removal
  cluster_size_histogram: {"2": 15_100, "3-10": 7_600, "11-100": 1_100, "101+": 100}
  largest_cluster_note: "auto-acknowledgement template; flagged as boilerplate"
  labeled_pair_eval: {pairs: 400, precision: 0.94, recall: 0.88}
  config_hash: 7f3c...

Define the rate explicitly: records removed divided by records in, computed per stage. Report the near-duplicate rate after exact removal, otherwise exact copies are double counted. The histogram tells a reviewer whether duplication is many small pairs (often harmless re-sends) or a few huge clusters (templates and boilerplate that need a filter rather than dedup).

Keep one, down-weight or keep all

Collapsing each cluster to one representative is the default for pre-training, but it is not always right for licensed operational data. Which record survives matters: keep the most complete, the most recent, or the one with the richest metadata, and write that rule down.

SituationSuggested handling
Exact re-sends, export artifactsKeep one; drop the rest
Document versions (draft, final, signed)Keep the final; optionally keep diffs for version-aware tasks; see RAG corpus duplicates and versions
Agent macro replies inside distinct ticketsKeep tickets; down-weight or mask the repeated span
Near-duplicates across train and eval splitsRemove from eval, or split by cluster ID
Large template clustersFilter as boilerplate, then dedup the rest

Split train, validation and test by cluster ID, not by record ID, so near-duplicates cannot leak across splits [1]. Lexical MinHash misses paraphrases and translated copies; for those, use semantic deduplication with embeddings, and for overlap with corpora you already hold, run a cross-dataset overlap check.

What to ask a data supplier before delivery

Ask the supplier how records were exported, whether duplicates or versions were removed upstream, and what fields identify the source record, because dedup decisions depend on that context. Useful requests include a stable record ID, source system, created and modified timestamps, thread or document IDs, and version markers, as covered in metadata fields for licensed text corpora. Then run your own exact and near-duplicate pass on a sample before full acceptance, alongside the checks in evaluating data supplier quality.

Note that de-identification can change duplication: if names and account numbers are replaced with consistent placeholders, records that differed only by customer become exact duplicates. Run dedup after redaction as well as before, and compare the rates. SourceX removes or replaces personal details such as names, emails, phones and account numbers before delivery and records the method, so plan your dedup pass on the delivered, de-identified text. If you are scoping a request for operational text, you can describe the dataset you need to SourceX.

Source licensed text data for your training pipeline

SourceX sources operational datasets, such as support and sales histories, engineering records and documents, from US companies on request, and every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery. Categories are not inventory, and a request does not guarantee a match. Describe the text data you need.

Sources

  1. Lee, Ippolito, Nystrom, Zhang, Eck, Callison-Burch, Carlini (arXiv; ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499
  2. SABR Research, "Deduplication at Scale: MinHash and LSH From Scratch". https://sabrresearch.com/cookbooks/pure-python-minhash-lsh-deduplication
  3. CodeSignal, "Introduction to Dataset Deduplication". https://codesignal.com/learn/courses/optimized-data-preparation-for-large-scale-llms/lessons/dataset-deduplication-and-redundancy-removal
  4. PyData London 2026 (pretalx), "PyData London 2026" (2026). https://pretalx.com/pydata-london-2026/speaker/RUXK7W/
  5. Zilliz (Milvus documentation), "MinHash LSH". https://docs.zilliz.com/docs/minhash-lsh

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data