Skip to content

Data quality, coverage and contamination

Measuring Dataset Diversity: Lexical, Semantic and Task Diversity Metrics

Quick answer

Measure dataset diversity on three separate axes, because each catches a different kind of redundancy. Lexical metrics (distinct-n, compression ratio, Self-BLEU) show repeated wording and templates. Semantic metrics (embedding dispersion, cluster counts, Vendi Score) show whether different wording covers the same meaning. Task metrics count the intents, workflows and outcomes a dataset actually covers. Report all three on the same sample size, against a reference set, and read them next to quality scores so variety from noise is not mistaken for useful coverage.

By SourceX Editorial · Updated

Recent data-centric work lists diversity among metric-based quality checks next to duplication, fluency and factual accuracy [1], and it answers a different question from the rest of the training data quality assessment cluster. Near-duplicate detection with MinHash and LSH removes copies. Coverage gap analysis maps a dataset to your deployment distribution. Diversity measures how spread out the remaining records are, which is what tells you whether a second candidate dataset adds anything you do not already have.

Lexical diversity: what n-gram and compression metrics actually catch

Lexical metrics are cheap, reference-free and good at exposing templated or macro-generated text, but they are blind to paraphrase. Run them first, on every candidate, before spending GPU time on embeddings.

  • Distinct-n: unique n-grams divided by total n-grams, usually for n = 1 to 4. It drops as the corpus grows, so only compare values computed on equal token counts.
  • Type-token ratio (TTR): unique word types over tokens. Raw TTR is strongly length-dependent; use a moving-average or fixed-window variant (MATTR, MTLD) when record lengths differ between candidates.
  • Compression ratio: original bytes over gzip-compressed bytes for the concatenated sample. It is fast, scales to billions of tokens and tracks repeated n-gram structure well, which makes it a good first screen.
  • Self-BLEU and long n-gram self-repetition: how much each record overlaps with the others. Report these together with compression ratio, since each catches a different kind of repetition and one number alone is easy to game.

The common failure mode is support and sales data. Agent macros, signature blocks, legal footers and auto-acknowledgment emails inflate compression ratio and Self-BLEU while the human-written turns may be quite varied. Strip or tag those turns first, as described in separating bot, macro and human turns in conversation data, and compute lexical metrics on human text only.

Semantic diversity with embeddings: dispersion, clusters and the Vendi Score

Semantic metrics tell you whether records differ in meaning, and they depend entirely on the embedding model you choose. Embed each record (or each user turn) with one fixed model, normalize vectors, and compute several statistics rather than one; the embeddings glossary entry covers the basics.

  • Mean pairwise cosine distance: simple dispersion. It is dominated by the bulk of the data and hides small, tight clusters of near-paraphrases.
  • Cluster count and entropy: run k-means or HDBSCAN, then report the number of non-trivial clusters and the Shannon entropy of the cluster-size distribution. A dataset where 70% of records fall into a handful of clusters is narrow even if its pairwise distance looks healthy.
  • Vendi Score: the exponential of the Shannon entropy of the eigenvalues of a similarity matrix built from a similarity function you define. It reads as an effective number of distinct items and needs no reference distribution. Because it weights items by prevalence, a long tail of rare items barely moves it, so pair it with cluster counts or a rare-cell check.
  • Nearest-neighbor distance to your existing data: for each candidate record, the cosine distance to its closest neighbor in your current training set. This is the metric that most directly answers "is this purchase redundant".

Pin the embedding model name and version, the text field embedded, truncation length and the similarity threshold in the report, or results will not be comparable across vendors. When a supplier computes embeddings for you, remember that embeddings are not anonymous: Morris et al. recovered 92% of 32-token inputs exactly from their embeddings [4], so treat shared vectors with the same controls as the underlying text. The semantic deduplication guide uses the same vectors, so compute both in one pass.

Task and instruction diversity in SFT data

Task diversity is the axis buyers of fine-tuning data care about most, and it cannot be read off text statistics. It counts what work the dataset represents: which intents, which workflows, which outcomes, and how evenly they are spread.

For instruction and conversation data, build a taxonomy before you look at counts. Typical dimensions are intent (refund, troubleshooting, contract redline), workflow step (triage, diagnosis, resolution, escalation), outcome (resolved, escalated, churned, reopened), artifact type (email, ticket, call transcript, form) and difficulty (turn count, tools touched, handoffs). Tag a stratified sample with an LLM classifier plus human spot checks, then report coverage (share of taxonomy cells with at least k examples) and evenness (normalized entropy across cells).

LIMA showed that a 65B model fine-tuned on 1,000 carefully curated, deliberately varied prompt-response pairs could perform competitively, while noting that this curation is labor-intensive [2]. The practical reading for buyers is that task spread per record matters more than raw volume for SFT, which is also why instruction-tuning data quality filtering and diversity selection are usually run together. Rare cells deserve their own targets; see long-tail and edge-case coverage.

The diversity vs quality trade-off

Diversity and quality pull against each other at the margin, so never optimize a diversity score alone. OCR garbage, mis-encoded characters, mixed languages, truncated records and spam all raise distinct-n and embedding dispersion. Aggressive quality or toxicity filters, in turn, tend to remove unusual dialects, informal registers and rare domains, which lowers diversity.

The defensible approach is to fix a quality floor first, then measure diversity only on records that pass it. Report diversity before and after each filter so you can see what the filter removed; the pretraining text quality filtering guide and toxicity filtering trade-offs describe the side effects to watch. For synthetic data, some recent evaluation frameworks such as SynEval add diversity as its own dimension alongside fidelity, utility and privacy [5], and generated sets often score well on lexical metrics while collapsing semantically; see synthetic data quality assessment.

A diversity scorecard you can send to a supplier

A comparable scorecard needs fixed sample sizes, named tools and a reference set, otherwise two suppliers' numbers cannot be compared. ISO/IEC 5259-2 provides a data quality model and measurable characteristics for reporting ML data quality [3], but its measures are general-purpose rather than text-specific, so specify the exact computations yourself.

Illustrative example: invented to show structure; it does not describe an available dataset.

AxisMetricComputation specCandidate ACandidate BYour current setWhat it tells you
LexicalDistinct-22M-token random sample, lowercase, whitespace tokens0.410.290.38B reuses phrasing heavily
LexicalCompression ratiogzip on concatenated sample2.94.13.0B is template-heavy
LexicalSelf-BLEU-41,000 records vs 1,000 references0.220.470.25Confirms B templating
SemanticVendi Score (cosine)5,000 records, pinned embedding model410160380A has more effective distinct items
SemanticMedian NN distance to current setSame model, cosine0.310.12n/aB mostly overlaps what you own
TaskTaxonomy cells with 20+ records60-cell intent x outcome grid442130A fills 14 cells you lack
TaskNormalized cell entropyOver filled cells0.830.550.71B is concentrated in a few intents
Quality floorRecords passing filterLanguage ID, length, encoding checks93%97%95%Diversity measured after this filter

Send the computation-spec column with your request so the supplier runs identical settings, and ask for the raw per-record outputs (cluster IDs, nearest-neighbor distances, taxonomy tags) rather than summary numbers alone. The dataset quality report contents guide explains how to fold these into a full report.

Common measurement mistakes

Most disagreements about diversity come from measurement choices, not data. Watch for these:

  1. Comparing distinct-n or TTR across samples of different sizes.
  2. Measuring before removing exact and near-duplicates, so the same metric double-counts redundancy.
  3. Switching embedding models between candidates, or embedding whole threads in one candidate and single turns in another.
  4. Treating the Vendi Score as a coverage measure; rare items barely move it and it says nothing about your target distribution.
  5. Counting task labels a classifier assigned without auditing a sample; see annotation quality audits.
  6. Reporting one aggregate number for a multi-source dataset instead of per-source and per-period breakdowns.

Finding varied operational data for your gaps

Once your scorecard shows which intents, workflows or outcomes you lack, the next step is a precise data description. SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents and finance and legal workflows, and manages licensing and ongoing purchases. Nothing is held in stock, so you describe the data you need, not the businesses, and a request does not guarantee a match. You can describe the coverage you need to SourceX using the taxonomy cells from your scorecard.

Describe the varied data you need

Every SourceX dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and personal details are removed or replaced before delivery. Diligence materials are prepared per dataset, and nothing is contracted until a supplier agrees. Share your diversity targets and missing task cells at sourcex.si/buyers.

Sources

  1. arXiv, "Towards Next-Generation LLM Training: From the Data-Centric Perspective" (2026). https://arxiv.org/pdf/2603.14712
  2. arXiv (Zhou et al.), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  3. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
  4. arXiv (Morris et al.), "Text Embeddings Reveal (Almost) As Much As Text" (2023). https://arxiv.org/pdf/2310.06816
  5. IEEE Computer Society Technical Committee on Security and Privacy, "IEEE S&P 2026 poster: SynEval synthetic data evaluation" (2026). https://www.ieee-security.org/TC/SP2026/downloads/posters/sp2026posters-final90.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data