Skip to content

Tables, time series and transactional data

Target Leakage in Licensed Tabular Data: How Buyers Audit Before Training

Quick answer

Target leakage in tabular data happens when a column carries information that only exists after the outcome you want to predict, so a model scores well in validation and fails in production. In supplier tables it usually hides in status fields updated after the event, the supplier's own risk scores, time-encoding IDs and recomputed totals. Buyers catch it before purchase by mapping the prediction point, requesting column-level population timestamps, and running single-feature baselines on the sample as an acceptance test.

By SourceX Editorial · Updated

What counts as target leakage, and what does not

Target leakage is a feature problem: a column encodes the label because it was written at or after the target event. Practitioner guides separate it from two neighbors, temporal leakage (training on rows from the future) and pipeline contamination (fitting scalers, imputers or selectors before the split) [1]. All three produce the same symptom, a validation score that does not survive contact with unseen data [1].

The distinction matters for who fixes it. Pipeline contamination is your own code: scalers, imputers and feature selectors must be fit inside the training folds, which in scikit-learn means wrapping StandardScaler, SimpleImputer and selectors in a Pipeline [1]. Target leakage, by contrast, is baked into the table you license, and no amount of careful modeling removes a column whose meaning is "the outcome, restated."

This page covers columns that encode the outcome. Train/test overlap from duplicate entities, and evaluation-set exposure, are different failure modes; see benchmark contamination, contamination checks for licensed evaluation data and, for model-side exposure of records, whether models memorize and leak data.

Why licensed operational tables leak more than curated benchmarks

Operational tables leak because they are snapshots of systems that keep writing after the event. A CRM opportunities export pulled today carries the current stage, close_reason and last_activity_date, not the values as of the day a forecast would have been made. The table looks like one row per entity, but each column was populated on its own clock.

The problem is well documented in applied work. In predictive process monitoring, for example, researchers studying event-log benchmarks have pointed to inconsistent splits and leakage across cases as a reason results are hard to reproduce [6]. The same mechanics apply to any table exported from a live workflow system.

Supplier-side patterns worth checking first (treat these as review hypotheses, not a complete list):

  • Post-outcome status fields. account_status = 'closed_charged_off' in a credit table, churn_flag_reason in a subscription table, claim_disposition in a claims extract.
  • The supplier's own model output. A fraud_score or risk_tier produced by a model that was itself trained on the outcome, or refreshed after a chargeback.
  • IDs and sequence numbers that encode time. Monotonic invoice_no, ticket numbers or Snowflake-style IDs let a model infer recency, and sometimes which rows were created by a recovery or collections workflow.
  • Recomputed aggregates. lifetime_value, total_paid or balance recalculated after refunds, write-offs or adjustments that were triggered by the outcome.
  • Process artifacts. A non-null collections_agent_id, escalation_queue or manual_review_at that only exists because the bad outcome occurred.
  • Missingness as signal. Fields that are null for open cases and filled for resolved ones, so is_null(resolution_date) predicts the label.

Map the prediction point before you look at any column

The audit starts with three timestamps, not with the data: the prediction point (when the model will score a row in production), the target event (when the label becomes true), and the allowed feature window (what may be known at the prediction point) [2]. Write these down per use case, because the same table can be clean for one task and leaky for another.

A churn model scoring accounts on the first of each month may use usage through the last day of the prior month. A fraud model scoring at authorization may use nothing written after the authorization timestamp, which rules out every dispute, chargeback and manual-review field. A credit model at origination may not use any payment-history column at all.

Feature stores formalize this as a point-in-time join. In Feast, the entity dataframe's event_timestamp is an inclusive upper bound, and retrieval returns the latest feature value stamped at or before it [4]. If a supplier cannot give you per-column timestamps, you cannot reproduce that join, and you are trusting a snapshot.

For time-ordered tables, a chronological split is a necessary companion check, because a random shuffle trains on the future [3]. How to construct those splits belongs in your evaluation design; this page is about which columns are allowed to exist in the first place. For table grain, keys and history requirements, see what to specify when licensing tabular data.

What to request in the data dictionary

The most useful leakage control is contractual documentation, requested before the sample. Ask the supplier to extend the data dictionary with fields that let you place every column on the timeline. A plain list of column names and types is not enough.

Illustrative example: invented to show structure; it does not describe an available dataset.

Dictionary fieldWhat it answersLeakage it exposes
column_name, source_system, source_tableWhere the value originates (ERP, CRM, billing, ticketing)Columns copied from a downstream collections or claims system
populated_at_eventWhich business event writes the value (created, updated, resolved)Fields written only at resolution
last_modified_semanticsOverwritten in place, append-only, or versioned (SCD Type 2)Current-state values masquerading as historical
as_of_availableWhether a historical value can be reconstructed for any past dateSnapshot-only columns that cannot be point-in-time joined
derivationFormula or upstream model for computed fieldsSupplier model scores and post-adjustment totals
null_meaningWhy the field is null (not applicable, not yet known, not captured)Missingness that tracks case status
target_relationSupplier's own statement of whether the column is affected by the outcomeColumns the supplier already knows are downstream

The ISO/IEC 5259 series gives a vocabulary for this conversation; Part 4 sets out a data quality process framework for training and evaluation data in supervised ML [5]. Buyers do not need to cite it in a request, but it helps when procurement asks why a dictionary review is a gating step. Related grounding on dictionaries themselves is in data dictionaries and schema documentation for AI.

Run leakage tests on the sample

Leakage tests on a sample are cheap, statistical and decisive enough to block a purchase. Run them on the supplier's sample before any payment milestone, using a frozen script you share in advance so results are not disputed later.

Illustrative example: invented to show structure; it does not describe an available dataset.

Pre-purchase leakage acceptance checklist

  1. Single-feature baselines. Fit a shallow model (depth-2 tree or logistic regression) on each column alone against the label. Any single feature with AUC near 1, or far above a domain-plausible ceiling, goes to manual review.
  2. Importance concentration. Train a gradient-boosted model on all columns. If one or two features carry most of the gain, inspect their populated_at_event first.
  3. Null-pattern test. Replace each column with is_null(column) and rerun the single-feature test. A strong missingness signal usually means resolution-time population.
  4. ID and sequence test. Score the label on rank(id) and on row order. Predictive IDs point to time encoding or batch exports split by outcome.
  5. Timestamp ordering. For every column with a timestamp, count rows where the value's timestamp is after the target event. The tolerated count is zero for anything in the feature set.
  6. Chronological holdout gap. Compare a random split against a time-ordered split [3]. A large drop on the time split flags temporal leakage or drift that the supplier should explain.
  7. Pipeline hygiene on your side. Confirm every fitted transform runs inside a Pipeline within cross-validation [1], so you are not blaming the supplier for your own contamination.
  8. Disposition log. Record each flagged column as drop, keep, or reconstruct as-of, with the reason, so the license and delivery spec can reference the agreed feature list.

A suspicious baseline is a trigger for review, not proof of leakage. Some outcomes really are easy to predict, and some strong features are legitimate. The point is to force an explanation before training budget is spent.

Worked example: a churn table that looked too good

Illustrative example: invented to show structure; it does not describe an available dataset.

A buyer evaluates a subscription table with 40 columns and a churned_90d label. The single-feature test returns an AUC near 1 for cancellation_survey_score, high for last_login_days, and surprisingly high for support_tier.

Mapping the timeline resolves each one. The survey is shown during the cancellation flow, so it is post-outcome and is dropped. last_login_days is computed at export time, so the supplier is asked for an as-of reconstruction from login events. support_tier is downgraded by billing automation after a missed payment that often precedes churn; whether it stays depends on whether the downgrade happens before the prediction point, which requires its last_modified_semantics.

The result is a smaller, honest feature set and a delivery spec that names the reconstruction. That is the same discipline that applies to labels themselves; see verifying outcome labels in operational records.

Leakage in tabular foundation model and fine-tuning data

Leakage matters differently when tables feed pretraining or LLM fine-tuning rather than one supervised model. For tabular foundation model pretraining, a leaky column in a held-out evaluation table inflates benchmark results in the same way it inflates a churn AUC. For table-reasoning fine-tuning, post-outcome columns can teach a model that the answer is usually sitting in another field.

In both cases, keep the evaluation tables to the same standard as supervised training data: documented prediction points, timestamped columns and a recorded disposition for every flagged feature. A short written record per feature, stating when it is populated and why it is allowed, is the artifact reviewers will ask for later.

How SourceX approaches tabular requests

SourceX sources operational datasets, including support and sales histories, engineering records and finance workflows, from US companies on request; categories are not inventory, and a request does not guarantee a match. During Assess, SourceX reviews the data and the licensing permissions, and diligence materials covering source, rights, preparation and allowed use are prepared per dataset. Every dataset is delivered under a license that defines records, uses, term and delivery, so the agreed feature list and as-of reconstructions you specify can be written into the request. Buyers can describe the table, label and prediction point they need, and SourceX looks for US businesses that hold it.

For the wider cluster, start at the tabular, time-series and transactional data buyer's guide or the AI data hub.

Request tabular data with documented outcome timing

If your model depends on knowing when each column was written, put the prediction point and feature window in the request. SourceX manages the process from Find through Manage, and nothing is contracted until a supplier agrees. Describe the data you need.

Sources

  1. Machine Learning Mastery, "3 subtle ways data leakage can ruin your models and how to prevent it". https://machinelearningmastery.com/3-subtle-ways-data-leakage-can-ruin-your-models-and-how-to-prevent-it/
  2. sharedcontext.ai, "leakage-guard". https://sharedcontext.ai/skills/external/zpower426/leakage-guard
  3. temporalcv documentation, "Why Time Series Is Different". https://temporalcv.readthedocs.io/en/latest/guide/why_time_series_is_different.html
  4. Feast documentation, "Feature retrieval". https://docs.feast.dev/v0.63-branch/getting-started/concepts/feature-retrieval
  5. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Data quality for analytics and ML, Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
  6. Weytjens and De Weerdt, arXiv, "Creating unbiased public benchmark datasets with data leakage prevention for predictive process monitoring" (2021). https://export.arxiv.org/abs/2107.01905

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data