Data quality, coverage and contamination
Semantic Deduplication with Embeddings: When MinHash Is Not Enough
Quick answer
Semantic deduplication embeds every record, groups records whose vectors sit within a cosine-similarity threshold, and keeps one representative per group. It catches paraphrases, templated rewrites and translated restatements that MinHash misses because they share meaning but few shingles. Add it after lexical dedup when your corpus mixes sources that restate each other, such as support macros, SFT pairs or synthetic rewrites. Calibrate the threshold on hand-labeled pairs, because a single global cutoff deletes legitimately distinct cases.
By SourceX Editorial · Updated
What semantic dedup catches that MinHash and LSH miss
Semantic dedup catches duplicates of meaning rather than duplicates of surface text. MinHash estimates Jaccard similarity over character or word n-gram shingles, so two documents must share a large fraction of literal shingles to collide in LSH buckets; that is exactly the property Lee et al. relied on to remove near-verbatim repetition, such as a 61-word sentence repeated over 60,000 times in C4 [1]. A paraphrase that swaps vocabulary and word order can drop Jaccard similarity far below any practical MinHash threshold while remaining the same training signal [3]. For the lexical layer itself, see our guide to exact and near-duplicate detection with MinHash and LSH.
The failure cases that slip past lexical dedup in buyer-sourced text are predictable:
- Support macros with agent edits. The same canned reply, personalized in every ticket, produces thousands of semantically identical assistant turns.
- SFT pairs regenerated from one seed. A prompt rewritten five ways by an LLM, each with a near-identical answer, overweights one instruction.
- Knowledge-base articles and their copies. A help-center article pasted into tickets, wiki pages and email threads in slightly different wording.
- Translated or summarized restatements. A multilingual embedding model maps a Spanish and English version of the same policy close together; shingles never will.
- Boilerplate-heavy records. Contracts or incident reports where the distinctive content is a small fraction of the text (this one cuts the other way, as discussed below).
How embedding-based deduplication works
Embedding-based dedup is a nearest-neighbor problem: embed, cluster, compare within clusters, then choose survivors. Embedding-based methods, SemDeDup among them, compare vectors from a pretrained model to find records that are semantically close but not lexically identical, then remove the redundant ones [3]. In practice the pipeline has five steps.
- Normalize and chunk. Strip signatures, ticket IDs and quoted reply chains first, or they dominate the vector. Decide the unit: whole conversation, single turn, or prompt-response pair. Embedding models truncate at their context limit, so long documents need chunking or a summary field.
- Embed. Use a sentence-embedding model suited to the domain and language mix (see the embeddings glossary entry). L2-normalize vectors so inner product equals cosine similarity.
- Partition. Run k-means (or reuse lexical cluster IDs) so pairwise comparison happens inside clusters, not across the whole corpus.
- Compare and group. Within each partition, use an approximate nearest-neighbor index (FAISS IVF or HNSW, or a vector database) to find pairs above the threshold, then take connected components or a greedy pass to form duplicate groups.
- Select survivors. Keep one record per group by a stated rule: highest quality score, most recent, longest resolved thread, or the human-written original over a synthetic rewrite. Log every dropped ID with its survivor ID and similarity so the decision is reversible.
The payoff depends on how much restatement your sources contain. Published removal rates come from web-scale corpora, not from licensed domain text, so treat them as evidence the method is worth testing rather than a removal rate to expect; measure your own rate on a held-out slice before committing compute.
Choosing a cosine similarity threshold for dedup
There is no standard cosine similarity threshold for dedup; the right value depends on the embedding model, the record unit and how much near-variation your task needs. As one documented example, the CuratorKIT post-training pipeline combines MinHash with an embedding dedup pass at 0.92 [2]. Treat that as a starting point to calibrate, not a default to copy, because cosine scores are not comparable across models: one model's 0.92 can be another's 0.85.
Calibrate with a labeled pair sample. Draw candidate pairs stratified across similarity bands (for example 0.80-0.85, 0.85-0.90, 0.90-0.95, above 0.95), have two reviewers label each pair, and pick the threshold where precision of "duplicate" calls meets your tolerance. Track inter-rater agreement on the labels; if reviewers disagree on what counts as a duplicate, the threshold cannot fix that (see inter-annotator agreement metrics).
Illustrative example: invented to show structure; it does not describe an available dataset.
| Similarity band | Pairs labeled | Labeled true duplicate | Labeled distinct | Typical distinct-case pattern | Decision |
|---|---|---|---|---|---|
| above 0.95 | 200 | 196 | 4 | Same macro, different order number | Drop automatically |
| 0.92-0.95 | 200 | 171 | 29 | Same symptom, different root cause | Drop only if outcome fields match |
| 0.88-0.92 | 200 | 104 | 96 | Same product, different issue | Keep; flag for diversity sampling |
| below 0.88 | 200 | 12 | 188 | Topically related | Keep |
The table shows the useful shape of the decision: a hard cutoff where precision is high, a conditional band where a structured field breaks ties, and a review band you do not auto-delete.
Failure modes: when semantic dedup removes data you need
Aggressive semantic dedup deletes legitimately similar but distinct cases, and in operational data those are often the most valuable records. Two support tickets that describe the same error message can end with different resolutions, a refund versus a firmware fix; their embeddings are close because the customer text dominates, but the outcome is the label you are training on. The same holds for contract clauses that differ by one negation, or incident reports that differ only in the failure component.
Guard against this with field-aware rules:
- Condition on outcome fields. Only treat a pair as duplicate if structured fields such as resolution code, disposition or final status also match (see verifying outcome labels in operational records).
- Embed the distinctive part. For boilerplate-heavy documents, embed the non-template spans, or every contract from one template collapses into one group.
- Protect the long tail. Check that removal does not disproportionately hit rare categories; compare category histograms before and after, as in long-tail and edge-case coverage.
- Watch diversity, not just volume. Semantic dedup should raise embedding-space diversity; if your dataset diversity metrics do not move, the pass removed little real redundancy.
- Do not confuse dedup with contamination checks. Removing paraphrases of your eval set from training is a separate, stricter job with its own threshold and audit trail.
Controlling cost: lexical first, semantic within clusters
Run exact hashing and MinHash first, then run semantic dedup only on what survives, because lexical passes are cheap per record and remove the bulk of verbatim repetition before you pay for embeddings. Semantic dedup adds the cost of computing an embedding for every record plus the nearest-neighbor search over them [3]. Vendors working at trillion-token scale describe deduplication as a major engineering bottleneck in its own right [4], and that is a self-reported vendor view, but the arithmetic is simple: embedding compute scales with record count and length, and exhaustive pairwise comparison scales quadratically.
Practical levers, roughly in order of savings:
- Shrink the input. Exact hash, then MinHash at your usual Jaccard threshold, before embedding anything.
- Partition hard. Compare only within k-means clusters, language, product line or source system; cross-partition duplicates are rarer and can be sampled.
- Pick the right unit. Embedding a 40-turn conversation as one vector is cheaper and often better than embedding every turn, unless the turn is your training unit.
- Use a smaller model for recall, a bigger one to confirm. Screen with a fast encoder and re-score only candidate pairs above a loose cutoff.
- Store vectors as an artifact. Persist embeddings in Parquet with record ID, model name and version, so later overlap checks and new deliveries reuse them.
Running semantic dedup on licensed and purchased data
On purchased data, semantic dedup belongs in acceptance testing and in overlap checks, and the vectors themselves need the same controls as the text. Embeddings are not anonymized derivatives: Morris et al. show an iterative inversion method recovering 92% of 32-token inputs exactly from their embeddings [5]. Keep the vector store inside the same access boundary as the licensed corpus, and delete or retain it on the same schedule the license sets for the source records.
Three buyer-side uses matter most. First, deduplicate a new delivery against itself to measure real record count versus billed record count. Second, run the same embeddings against data you already hold to find paraphrased overlap, as in cross-dataset overlap checks. Third, test novelty against open corpora you already train on, covered in testing licensed data against public web crawls. Write the dedup unit, model, threshold and survivor rule into the acceptance criteria so a supplier and buyer measure the same thing; our data quality assessment hub lists the other checks that sit beside it.
If you are sourcing operational text such as support histories, sales conversations or engineering records, SourceX's buyer intake starts from a description of the data you need, not from a catalog.
Sourcing text data you can deduplicate with confidence
SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records and documents, and every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, with the method recorded and a sample checked. Describe the data you need at sourcex.si/buyers.
Frequently asked questions
Is semantic dedup a replacement for MinHash?
No. MinHash is cheap and precise for near-verbatim copies, and running it first shrinks the set you must embed. Semantic dedup is a second layer for paraphrases and restatements that lexical similarity cannot see [3].
Which embedding model should I use for deduplication?
Use one that matches your language mix and record length, and record its name and version with the vectors. Thresholds do not transfer between models, so changing the model means re-running threshold calibration on labeled pairs.
Should SFT pairs be deduplicated on the prompt, the response, or both?
Usually both, with different rules. Near-identical prompts with materially different responses can be useful preference or diversity signal, while near-identical prompt-response pairs mostly overweight one behavior; a concatenated embedding plus a response-only check separates the two cases.
Sources
- Lee et al. (arXiv), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499
- arXiv, "CuratorKIT: Data Curation and Synthetic Data Generation for LLM Post-Training" (2026). https://arxiv.org/pdf/2606.21631
- ZeroEntropy, "Deduplication". https://www.zeroentropy.dev/concepts/deduplication/
- Zilliz, "Data Deduplication at Trillion Scale: How to Solve the Biggest Bottleneck of LLM Training". https://zilliz.com/blog/data-deduplication-at-trillion-scale-solve-the-biggest-bottleneck-of-llm-training
- Morris, Kuleshov, Shmatikov, Rush (arXiv; EMNLP 2023), "Text Embeddings Reveal (Almost) As Much As Text" (2023). https://arxiv.org/pdf/2310.06816
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.