Data quality, coverage and contamination
Checking Overlap Between a New Dataset and Data You Already Own
Quick answer
A dataset overlap check measures what share of a candidate dataset's records already exist, exactly or nearly, in data you have licensed or collected. Index your existing holdings with exact hashes, MinHash signatures and embeddings, run the supplier's sample against that index, and report overlap as a percentage with confidence bounds. When neither side can share raw records, use hashed identifiers or private set intersection. Set a maximum acceptable overlap before the sample arrives, and tie price and acceptance to the result.
By SourceX Editorial · Updated
Why cross-holdings overlap is a purchase decision
Overlap with what you already own is a cost problem first: you pay again for records that add no new information, and the duplicates can quietly distort training and evaluation. Within-dataset deduplication is a cleaning task, covered in our guide to exact and near-duplicate detection with MinHash and LSH. Overlap with the open web is a novelty question, covered in testing licensed data against public web crawls. This page, part of our training data quality assessment hub, is about the third case, where the duplicate is already on your own storage.
That third case is common because operational data travels. The same support transcripts can reach you through a BPO vendor and through the client company; the same contracts can sit in a legal-tech vendor's corpus and in a document aggregator's. Two suppliers can also both resell a shared upstream source. If you only check the new dataset against itself, none of this shows up.
Hidden overlap is not hypothetical even in curated data. Lee et al. found that train/validation overlap in standard language datasets affected more than 4% of validation examples, and that deduplication reduced how often models emitted memorized text [1]. The same mechanism applies across purchases: records that appear in both a fine-tuning set and an evaluation set you built from another vendor inflate your scores.
How overlap inflates evaluation and wastes spend
Duplicated records inflate evaluation because the model is scored on examples it has already seen, so held-out accuracy partly measures memorization rather than generalization [1]. If you plan to carve an evaluation split from a new purchase, or to use a new dataset to test a model trained on older holdings, overlap directly biases the result. Our page on decontaminating a licensed training set against public benchmarks covers the benchmark version of this problem.
For pre-training and SFT, the cost is subtler. Repeated examples get extra effective weight, which can over-represent one customer's phrasing, one product line or one template. Lee et al. reported a single 61-word sentence repeated over 60,000 times in C4, which shows how far unmanaged repetition can go in large corpora [1].
The commercial effect is easy to state. If 30% of a candidate dataset already sits in your lake, the effective price per novel record is about 43% higher than the quoted price per record. That ratio belongs in the negotiation, not discovered after delivery.
Choosing a matching method for each data type
Pick the matching method by what "the same record" means for your data, then run several layers, because each catches a different kind of duplicate. Exact hashing catches byte-identical copies; MinHash over shingles catches lightly edited or re-templated text; embeddings catch paraphrases and translations; key-based joins catch the same business event exported in different formats.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Layer | What it matches | Typical implementation | Failure mode to watch |
|---|---|---|---|
| Normalized exact hash | Identical records after whitespace, case and Unicode NFC normalization | SHA-256 of normalized text or canonical JSON | Misses a ticket with a changed timestamp or signature line |
| Natural key join | Same business event in different exports | Hash of ticket ID + created date + account hash, invoice number + vendor ID, Git commit SHA | Suppliers rekey IDs; keys collide across tenants |
| MinHash + LSH | Near-duplicates with edits, boilerplate or template changes | 5-gram word shingles, 128 to 256 permutations, banded LSH, Jaccard threshold around 0.8 | Long shared boilerplate (legal footers, email disclaimers) creates false positives |
| Embedding similarity | Paraphrases, summaries, translated copies | Sentence or document embeddings in an ANN index, cosine threshold tuned on labeled pairs | Topical similarity mistaken for duplication |
| Media perceptual hash | Re-encoded audio, images or video frames | pHash or audio fingerprints per frame or segment | Crops and speed changes evade simple hashes |
Vector databases and libraries offer MinHash LSH indexes that you build once over your holdings and then query with new items, which suits a buy-side check where the index persists across many supplier evaluations [2]. Strip known boilerplate before shingling, or every email thread with the same legal disclaimer will look like a near-duplicate. Calibrate thresholds on a few hundred hand-labeled pairs from your own data rather than reusing a default.
Running the check on a supplier sample
Run the check as a fixed procedure: build the index on your holdings, query a representative sample, adjudicate borderline matches, and extrapolate with an interval. The sample matters as much as the method; if the supplier hand-picks it, overlap on the sample tells you little about the full set, which is the topic of checking whether a vendor's sample is representative.
- Index holdings once. Include every licensed dataset, internal corpus and evaluation set, each tagged with source, license ID and intended use. Store exact hashes, MinHash signatures and embeddings side by side.
- Get a random sample. Ask for records drawn uniformly or stratified by time period and source system, with the sampling method stated in writing.
- Apply the same normalization and redaction to both sides. If the supplier replaced names with tokens like
[PERSON_1], apply the same placeholder style to your holdings, or redaction differences will hide matches. - Query each layer and record the strongest match per record: exact, key, near-duplicate (with Jaccard score) or semantic (with cosine score).
- Adjudicate a slice by hand. Review 50 to 100 flagged pairs and 50 unflagged records to estimate precision and miss rate.
- Report overlap share with a confidence interval, broken down by matched holding, so you know which earlier license the duplicates came from.
For the record count behind step 5, see sample sizes for estimating a dataset's error rate. A Wilson interval on the adjudicated overlap rate is usually adequate for a go/no-go decision.
Privacy-safe overlap when raw data cannot leave either side
When neither party can share raw records before a contract, compare fingerprints instead of content. The simplest option is keyed hashing: both sides compute HMAC-SHA-256 over normalized records or natural keys with a key agreed for this check only, then compare digests. This works for exact and key matches but leaks membership of every matched record to the other side.
Private set intersection (PSI) reduces that leakage. PSI protocols, commonly built on Diffie-Hellman key exchange, let two parties compare encrypted sets so that only the intersection is learned, and cardinality variants (often called PSI-C) return only the size of the intersection rather than which items overlap. For a pre-purchase check, the size is often all you need: "4,120 of 50,000 sample keys are already in your holdings" supports a pricing conversation without revealing which customers or documents are involved.
PSI handles exact matches only. For near-duplicate checks under confidentiality, options include running MinHash signatures through PSI band by band, or having a neutral party run the comparison inside a clean room or a zero-copy share such as Delta Sharing, with only aggregate counts released. Agree in writing that the fingerprints are used only for the overlap check and deleted afterward.
Setting acceptance thresholds and contract terms
Decide the maximum acceptable overlap before you see results, write it into the pilot or purchase terms, and connect it to price or replacement. Buyer guides for data vendors recommend running a paid pilot on real data to verify quality claims before contract terms are finalized, rather than relying on proposals or demos [3]. ISO/IEC 5259-2 offers a vocabulary of measurable data quality characteristics that can anchor how you define and report these measures [4].
Illustrative example: invented to show structure; it does not describe an available dataset.
overlap_acceptance:
holdings_index_version: "lake-2026-09-30"
sample: { size: 5000, method: "uniform random, supplier attests in writing" }
match_layers: [exact_normalized, natural_key, minhash_j0.8, embedding_cos0.92]
thresholds:
max_overlap_total: 0.10 # all layers combined
max_overlap_eval_sets: 0.00 # no overlap with internal eval splits
max_overlap_single_holding: 0.05
on_breach:
- "remove matched records from delivery and reprice"
- "or decline without obligation"
fingerprint_handling: "HMAC key destroyed after check; digests deleted within agreed period"
Treat overlap with evaluation sets differently from overlap with training data. A few percent duplication against an old SFT corpus may be tolerable at a lower price; any overlap with a held-out evaluation set should be removed before delivery. Pricing structures that charge per record or per token make it straightforward to exclude matched records, as discussed in our comparison of AI data license pricing structures.
Questions to ask suppliers about upstream sources
Overlap often has a provenance explanation, so ask where records came from before running any hashes. Data Cards and similar documentation frameworks recommend recording upstream sources and collection methods, which is exactly the information you need to predict overlap [5].
- Which systems produced these records (for example Zendesk, Salesforce Service Cloud, Jira, SAP), and over what date range?
- Have the same records, or records from the same upstream company, been licensed through any other channel or reseller?
- Is any portion derived from a dataset you or another buyer can obtain elsewhere?
- Were records merged, summarized or translated from other corpora?
- What identifiers persist after de-identification, so a key-based match is still possible?
Answers here feed your provenance file; our guide to consent and notice records covers the rights side of the same questions. If you track many vendors, the operating model in managing many data suppliers helps keep a single holdings register, and how to run a data pilot with a supplier shows where an overlap test fits in a pilot.
Where SourceX fits in an overlap-aware purchase
SourceX sources operational datasets such as support and sales histories, engineering records, documents, and finance and legal workflows from US companies on request, rather than holding stock, so buyers can describe the data they lack instead of picking from a catalog that may repeat what they hold. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and diligence materials on source, rights and preparation are prepared per dataset, which supports the provenance questions above. You can describe the records you need, including the gaps in your current holdings, on the SourceX buyers page.
Find data that adds to what you already hold
SourceX looks for US businesses that hold the data you describe and manages the process from Find and Assess through Agree, Transact and Manage, with nothing contracted until a supplier agrees. A request does not guarantee a match, and every release is approved by the supplying company. Describe the gap in your holdings at sourcex.si/buyers.
Sources
- Lee et al. (arXiv:2107.06499), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499
- Zilliz documentation, "MinHash LSH". https://docs.zilliz.com/docs/minhash-lsh
- CloudPano, "Selecting the Best AI Training Data Provider: A Practical Buyer's Guide". https://www.cloudpano.com/blog/selecting-best-ai-training-data-provider-buyers-guide
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- Pushkarna, Zaldivar, Kjartansson (Google Research), FAccT 2022, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.