Data quality, coverage and contamination
Detecting Drift Between Licensed Historical Data and Your Production Traffic
Quick answer
Training data drift detection before purchase means comparing a sample of the licensed historical records against a recent sample of your production inputs, field by field and in embedding space, and then explaining where they differ. Use PSI or Kolmogorov-Smirnov tests on structured fields, embedding-distance and classifier two-sample tests on text, and separate input shift from label and concept change. The output should be a segment-level report that tells you whether to reweight, filter, or ask the supplier for different slices.
By SourceX Editorial · Updated
This guide sits inside the training data quality assessment hub. It is narrower than coverage gap analysis, which asks whether your deployment segments exist in the data at all; here the question is how far the shared segments have moved.
Why historical operational data drifts from what your model will see
Historical business records drift because the systems, products, customers and policies that produced them change, so a ticket archive from 2021 describes a different world than this week's queue. A help desk migrates from Zendesk macros to an AI-assisted reply tool, a product line is retired, a new payment rail appears, or a form adds a required field. Each change alters field distributions, vocabulary and record length without anyone labeling it as a change.
The cost is measurable. The WILDS benchmark collected real-world shifts across domains and found substantial gaps between in-distribution and out-of-distribution performance, even for models that look strong on held-out data [1]. Document classifiers trained on one collection of business documents lose accuracy on out-of-distribution documents [6], and industrial sound detectors fail when operating conditions change the feature distribution [7]. If you fine-tune on licensed records that no longer resemble production, your held-out score from the same archive will overstate deployed performance.
Data drift vs concept drift: three shift types to separate
Data drift and concept drift are different failures, and a single "drift score" hides which one you have. Separate them before choosing a fix.
- Covariate shift (input drift): P(x) changes while the mapping from input to outcome holds. Example: production tickets are 40% mobile-app issues while the licensed archive is mostly desktop-web issues, but a refund request still means the same thing.
- Label shift (prior shift): P(y) changes. Example: chargeback disputes were 3% of historical cases and are now 9% of the queue. Class priors in the archive will miscalibrate a classifier.
- Concept drift: P(y | x) changes, so the same input now deserves a different label. Example: a "priority 2" ticket meant a four-hour SLA under the old policy and next-business-day under the new one. Fraud detection on transaction records is a well-documented case where concept drift arrives together with class imbalance and delayed labels [4].
Input-distribution tests catch only the first two. Concept drift requires checking label definitions over time, which is covered in code-set revisions and process changes in multi-year datasets, and outcome labels shaped by past decisions, covered in historical decision bias in operational labels.
Comparing training and production distributions on structured fields
For structured fields, compare each column's distribution with a test suited to its type, then rank fields by effect size rather than p-value. With large samples almost every difference is statistically significant, so the question is magnitude and whether it matters for the task.
Population stability index for training data. PSI bins a field (deciles of the reference sample for numerics, categories for categoricals) and sums (p_prod − p_hist) × ln(p_prod / p_hist) over bins. Credit-risk practice often treats values under 0.1 as stable and above 0.25 as a major shift; these are conventions, not statistical thresholds. Add a small epsilon to empty bins, fix bin edges from one sample, and report the top-contributing bins, because a PSI of 0.3 driven by a new "null" category means something different from one driven by a tail.
Kolmogorov-Smirnov. The two-sample KS statistic is the maximum distance between the two empirical CDFs and tests whether both samples come from the same continuous distribution [5]. Use it on amounts, durations, counts and timestamps-of-day. It is sensitive near the median and weak in the tails, so pair it with quantile comparisons (p1, p50, p99) for fields like invoice amount where tails carry the risk.
Categoricals and schema. Use chi-square or Jensen-Shannon divergence on category frequencies, and list categories present in production but absent from the archive: new product SKUs, new country codes, new status values. Schema-level differences (renamed columns, unit changes, enum remapping) should be resolved first using structured dataset validation checks, or every test downstream reports false drift.
Embedding drift detection for text and conversations
For text, embed both samples with the same frozen encoder and compare the resulting distributions, because token-level statistics miss shifts in topic and intent. Use a sentence-embedding model you will not fine-tune during the analysis, normalize the vectors, and strip boilerplate signatures and templated disclaimers first, since templates and canned replies can dominate distances.
Three complementary measures work well:
- Distance between distributions. Maximum mean discrepancy (MMD) with an RBF kernel, or Fréchet distance between fitted Gaussians, with a permutation test to get a null distribution from shuffled labels.
- Classifier two-sample test. Train a simple domain classifier (logistic regression on embeddings, or gradient boosting on embeddings plus metadata) to predict "archive vs production." A cross-validated AUC near 0.5 means the samples are hard to tell apart; an AUC near 1.0 means they are clearly different. The classifier's probabilities also give you per-record "production-likeness" scores you can reuse for reweighting.
- Cluster occupancy. Cluster the pooled embeddings (k-means or HDBSCAN), then compare the share of each cluster in archive and production. Clusters that exist only in production are coverage gaps; clusters that exist only in the archive are retired topics.
Cluster occupancy also separates two meanings of "representative" that are often conflated: an archive can mirror production proportions, or merely cover every production topic at some density, and the two call for different fixes [3]. For LLM pipelines, also check prompt length, language mix, attachment rates and the share of records containing structured artifacts such as JSON, stack traces or tables.
Explaining dataset differences, not just scoring them
A drift report is only useful if it says which segments moved and how, since a single MMD or AUC value cannot tell you whether to reweight or reject. Recent work frames distribution-shift explanation as finding interpretable transformations or feature-level descriptions that map one dataset toward another [2]. In practice, three outputs make a shift actionable:
- Top discriminating features from the domain classifier (SHAP values on metadata, or the highest-weight terms when you add a TF-IDF view).
- Representative examples from the most production-specific and most archive-specific clusters, read by a domain expert.
- Segment tables showing shares and outcome rates by channel, product, region, customer tier and year.
Run these on a sample large enough to estimate segment shares with useful precision; sample sizes for estimating dataset error rates gives the arithmetic. Deduplicate both samples first with MinHash and LSH near-duplicate detection, since repeated records inflate apparent differences.
A drift report you can run during a data pilot
Run the analysis on a supplier sample during evaluation, before full delivery, using the steps in how to run a data pilot with a supplier. Bring the supplier sample into your environment and run every comparison there, so production records never leave it.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Check | Field or view | Archive sample | Production sample | Statistic | Reading |
|---|---|---|---|---|---|
| Field drift | channel (email, chat, phone, app) | 62% email | 31% email, 38% app | PSI 0.41 | Major covariate shift |
| Field drift | first_response_minutes | p50 95, p99 2,880 | p50 12, p99 610 | KS 0.38 | Process change; feature not reusable |
| Category gap | product_line | 14 values | 17 values (3 new) | 9% of prod in unseen values | Coverage gap |
| Label shift | resolution_code = refund | 11% | 19% | Prior ratio 1.7 | Recalibrate or reweight |
| Text drift | Message embeddings | n = 5,000 | n = 5,000 | Domain AUC 0.86 | Clearly separable |
| Explanation | Top classifier features | n/a | n/a | channel, "app crash", "wallet" | Mobile-wallet issues missing from archive |
| Concept check | Priority definition | P2 = 4 h SLA | P2 = next business day | Policy doc diff | Do not reuse priority labels |
Store the report alongside the dataset version ID, the encoder name and version, bin edges and random seeds so the comparison can be rerun on later deliveries in Parquet or JSONL.
Deciding what to do: reweight, filter or request different slices
The right response depends on which shift you found and how much of production the archive covers. Use this decision table as a starting point.
| Finding | Typical response | Watch for |
|---|---|---|
| Moderate covariate shift, full support overlap | Importance-weight archive records by domain-classifier odds p/(1 − p), scaled by the archive-to-production sample-size ratio | Extreme weights; clip and check effective sample size |
| Archive covers production only in some segments | Filter to closer segments or years; keep the rest for pretraining-style mixing | Discarding rare but valuable edge cases |
| Production segments absent from archive | Request additional slices (recent years, channels, product lines) from the supplier | Small new slices that cannot be validated |
| Label shift only | Adjust class priors or resample to production rates | Unknown production priors without fresh labels |
| Concept drift in labels | Relabel a recent subset, or train on inputs only (for example, as SFT context) and drop stale labels | Mixing old and new label semantics in one model |
| Shift too large on all views | Use the archive for evaluation of robustness, not as primary training data | Paying for volume that adds little |
For evaluation sets, the logic reverses: a deliberately shifted held-out slice is useful for measuring robustness, as WILDS does with its out-of-distribution splits [1]. Keep a production-like test set as well, or you will not know which number to trust.
Specifying drift requirements in a data request
You can reduce drift before any data moves by describing the production distribution you need, rather than a generic category. State the date range that reflects current policy, the channels and product lines that matter, expected shares where you know them, and fields whose definitions changed. Ask suppliers for field dictionaries with effective dates, a record-count breakdown by year and segment, and a sample drawn across the full date range rather than the most convenient month.
SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents, and finance and legal workflows; categories are not inventory, and a request does not guarantee a match. Buyers describe the data they need, not the businesses, and every release is approved by the supplying company. You can describe a production-matched data request with the segment and date requirements from your drift report. Personal details are removed or replaced before delivery, which can itself change text statistics; see the de-identified data guide and run drift checks on the de-identified form you will actually receive.
Get historical data that matches your production traffic
SourceX finds US businesses that hold the operational data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Every dataset is rights-reviewed and delivered through private, access-controlled workflows after an executed agreement and supplier approval. Start with the segments, date ranges and fields your drift analysis flagged at sourcex.si/buyers.
Sources
- arXiv (Koh et al.), "WILDS: A Benchmark of in-the-Wild Distribution Shifts" (2021). https://arxiv.org/pdf/2012.07421
- Journal of Machine Learning Research, vol. 26 (Babbar, Guo, Rudin), "'What is Different Between These Datasets?' A Framework for Explaining Data Distribution Shifts" (2025). https://jmlr.org/beta/papers/v26/24-0352.html
- arXiv (Chasalow and Levy), "Representativeness in Statistics, Politics, and Machine Learning" (2021). https://arxiv.org/pdf/2101.03827
- IEEE TNNLS (Dal Pozzolo, Boracchi, Caelen, Alippi, Bontempi), "Credit Card Fraud Detection: A Realistic Modeling and a Novel Learning Strategy" (2018). https://boracchi.faculty.polimi.it/docs/2017_FraudsTNNLS.pdf
- NIST Information Technology Laboratory, "Dataplot Reference Manual: KS 2 Sample Test". https://itl.nist.gov/div898/software/dataplot/refman1/auxillar/ks2samp.htm
- NeurIPS 2022 Datasets and Benchmarks (Larson et al.), "Evaluating Out-of-Distribution Performance on Document Image Classifiers" (2022). https://proceedings.neurips.cc/paper_files/paper/2022/hash/4c0986bd04d747745beba3752bdf4d9d-Abstract.html
- arXiv (Tanabe et al., Hitachi), "MIMII DUE: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection with Domain Shifts" (2021). https://arxiv.org/pdf/2105.02702
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.