Skip to content

Tables, time series and transactional data

Subscription, Billing and Usage Histories for Churn Prediction Models

Quick answer

Churn prediction training data is a linked, multi-period set of account, subscription, invoice, payment, product-usage and support tables in which each customer's renewal or cancellation was actually observed. Widely used public churn CSVs are typically single snapshots with a pre-computed label, so they cannot teach a model when risk appears. To license useful data, fix the churn definition and prediction horizon first, then ask for complete cohorts, event timestamps, plan-change history and per-cohort base rates, and evaluate with customer-grouped, time-based splits.

By SourceX Editorial · Updated

What counts as churn has to be decided before you request data

The churn definition determines which tables and timestamps a supplier must deliver, so write it down before any sourcing conversation. B2B subscription businesses usually mean one of three things: contract non-renewal at term end, voluntary cancellation mid-term, or involuntary churn after failed payment and dunning. Downgrades (seat or tier reductions) and gross-revenue contraction are separate targets, and "inactivity churn" (no logins for N days while still paying) is a leading indicator rather than a revenue event.

Each definition needs different evidence. Non-renewal needs contract start, end and renewal dates plus the renewal decision; involuntary churn needs payment attempts, decline codes and dunning steps; contraction needs line-item quantities on every invoice. Pair the definition with a horizon, for example "predict, at day 0 of the last 90 days of term, whether the account renews," because the horizon decides how far back feature windows must reach and how much history per account you need.

Write the reactivation rule too. An account that cancels and resubscribes within 30 days may be one continuous customer or one churn plus one new logo; the supplier's billing system will not decide that for you.

Which tables a churn dataset needs and how they join

A trainable churn dataset is relational: five to six tables keyed by a stable account identifier, each with event-level timestamps. Flattening them into one row per customer before delivery destroys the temporal ordering you need for leakage-safe features, which is why research benchmarks such as RelBench frame churn-style tasks as entity-level predictions over linked tables [2]. RelBench v2 extends this to consumer-platform data, a sign that customer-behavior prediction over multi-table histories is now a standard research pattern [3].

The core record set:

  • Accounts: pseudonymous account ID, created date, segment, industry code, employee band, region, acquisition channel, parent-child hierarchy.
  • Subscriptions and plan changes: one row per subscription version with plan, seats or units, list and contracted price, term, start, end, change type (new, upgrade, downgrade, renewal, cancel) and cancellation reason code where captured.
  • Invoices and payments: invoice lines, amounts, currency, due and paid dates, payment method type, failed-attempt counts, decline reason, credit notes, refunds and write-offs.
  • Product usage events or daily aggregates: active users, sessions, feature-level event counts, API calls, storage, with the aggregation grain stated.
  • Support contacts: ticket created and closed timestamps, priority, category, CSAT, escalation flags; optional free text only if it has been de-identified.
  • Customer success touchpoints (if available): QBR dates, health-score history, and renewal opportunity stage, which overlaps with renewal workflow records.

Billing systems such as Stripe Billing, Zuora, Chargebee or Recurly, product analytics stores and helpdesk tools each keep their own IDs. Ask how the supplier joined them, what the match rate was and how unmatched rows were handled. For the general grain and keys checklist, see what to specify when licensing tabular data.

Illustrative request specification

A request spec lets suppliers say yes or no quickly and lets your reviewers compare offers on the same terms.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample specification
Churn eventNon-renewal at term end, plus voluntary mid-term cancellation; downgrades flagged separately
Prediction horizonScore at T-90 days before term end; label = renewed or not at T+30 grace
PopulationB2B accounts with annual or monthly contracts, all plan tiers, no segment exclusions unless disclosed
History depthAt least 24 months of observed billing per cohort, including accounts that churned and were archived
Tablesaccounts, subscription_versions, invoice_lines, payments, usage_daily, support_tickets
KeysStable pseudonymous account_id across all tables; subscription_id; invoice_id
Time fieldsUTC event timestamps; snapshot date for any slowly changing attribute
Base ratesChurn rate per signup cohort and per segment, as delivered
Personal dataNames, emails, phones, billing addresses and card fragments removed or replaced; method documented
FormatParquet or CSV per table with a data dictionary and row counts per file

Where churn models fail: leakage, survivorship and label drift

Most inflated churn scores come from three data defects you can check at intake: entity leakage, temporal leakage and survivorship bias. Random row-level splits put the same customer in both train and test, so the model memorizes accounts rather than learning risk; grouped, customer-level splits prevent this [1]. Time-based cut-offs, where features are computed only from data before the prediction timestamp and labels only after, are the second guard, and they mirror how relational benchmarks build train, validation and test windows [2].

Temporal leakage hides in columns. A status = cancelled field, a cancellation_reason, a final refund or a support ticket tagged "cancellation request" all encode the label; so does any attribute snapshot taken after term end. Ask the supplier which fields are point-in-time and which are current-state overwrites, because a CRM or billing system that updates plan_tier in place will silently leak the future.

Survivorship bias is the defect buyers most often miss. Retention and archival policies may have purged or anonymized churned accounts after a set period, which leaves a dataset of survivors and a churn rate far below reality. Ask for complete cohorts: every account that started in a given window, whatever happened to it.

Label drift matters for long histories. Pricing changes, a move from monthly to annual contracts, or a new dunning policy change the meaning of "churn" across periods, so request a dated changelog of plan catalog, billing policy and product releases alongside the tables.

Intake checks to run on a sample before you sign

A short intake on a sample answers most quality questions before money changes hands. Run these checks against whatever sample or schema the supplier can share:

  1. Cohort completeness: count accounts by signup month; a cliff in older cohorts suggests purged churners.
  2. Base-rate disclosure: compare your computed churn rate per cohort and segment with the supplier's stated rates; unexplained gaps usually mean exclusions.
  3. Key integrity: percent of invoice, usage and ticket rows that join to an account; orphan rates above a few percent need an explanation.
  4. Point-in-time audit: for each attribute, confirm whether history is versioned or overwritten.
  5. Usage grain: verify daily aggregates are complete for every active day and that zero-usage days are present, not just missing rows.
  6. Currency and tax: confirm amounts are net or gross of tax and whether multi-currency invoices are normalized.
  7. Censoring: identify accounts still active at the data cut-off; they are right-censored, not retained, and should be handled with survival methods or excluded from fixed-horizon labels.
  8. Identifier removal: spot-check free-text fields such as ticket bodies and invoice memos for names, emails and phone numbers.

If you also plan to use the same tables for revenue forecasting, the demand forecasting training data guide covers order-history specifics that differ from subscription data.

How suppliers can deliver multi-table billing and usage histories

Delivery format should preserve the relational structure: one file or table per entity, a data dictionary, and documented join keys. Parquet per table with explicit schemas is the simplest portable option. Warehouse-native sharing is another: Snowflake Secure Data Sharing exposes read-only objects to a consumer account without copying data [4], and Delta Sharing is an open protocol for reading shared tables across platforms [5].

Whatever the channel, the license should state which records are included, the allowed uses, the term and the delivery method. For ongoing purchases, such as monthly refreshes to retrain a health score, agree on how new cohorts and late-arriving events (back-dated credits, chargebacks) will be appended rather than overwritten.

Rights and privacy questions specific to customer billing data

Billing and usage records describe a supplier's own customers, so rights review centers on whether the supplier may license data about its customers and their end users. Customer contracts, the supplier's privacy notice and any data processing agreements can restrict secondary use of usage telemetry or support content. When the data is held by a service provider on behalf of clients, the authorization chain is different; see client data held by service providers.

Personal details still appear in B2B data: contact names on accounts, billing emails, user-level event logs and ticket text. Expect these to be removed or replaced, with the method documented, before you receive anything. Account-level aggregates can support many churn models without user-level events, which reduces exposure.

How churn data differs from CRM pipeline and renewal workflow records

Churn modeling needs post-sale financial and behavioral history, which is a different record set from pre-sale pipeline data. Opportunity stages, win and loss reasons and sales activities belong to sales CRM pipeline histories, which suit deal-scoring and forecasting rather than retention. Renewal processes themselves, including the tasks, approvals and communications around a renewal, are covered in customer renewal workflows and renewal workflow records; they add useful context features but rarely carry invoice-level or usage-level history on their own.

For the broader set of table, time-series and transaction data types, start from the structured data buyer's guide or the AI data hub. If your churn model is a relational deep learning model rather than gradient-boosted trees on engineered features, the multi-table relational data guide explains how to request whole schemas.

How SourceX handles requests for churn and retention data

SourceX sources operational datasets, including sales and support histories, finance workflows and documents, from US companies on request, and manages licensing and ongoing purchases. Nothing is held in stock and a request does not guarantee a match: you describe the data, SourceX looks for US businesses that hold it, and each release is approved by the supplying company. You can describe the churn and billing history you need using the specification above.

Each dataset is reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows only after an executed agreement. SourceX does not train models and does not publish prices; terms are agreed per deal.

License churn prediction training data through SourceX

If you need real subscription, invoice, usage and support histories with observed churn outcomes, SourceX can look for US companies that hold them and manage the process from assessment through license and delivery. Nothing is contracted until a supplier agrees. Describe the churn data you need.

Sources

  1. Machine Learning Mastery, "3 Subtle Ways Data Leakage Can Ruin Your Models (and How to Prevent It)". https://machinelearningmastery.com/3-subtle-ways-data-leakage-can-ruin-your-models-and-how-to-prevent-it/
  2. arXiv, "RelBench: A Benchmark for Deep Learning on Relational Databases" (2024). https://arxiv.org/pdf/2407.20060
  3. arXiv, "RelBench v2 (arXiv:2602.12606)" (2026). https://arxiv.org/html/2602.12606v1
  4. Snowflake Documentation, "About Secure Data Sharing". https://docs.snowflake.com/en/user-guide/data-sharing-intro.html
  5. Delta Lake documentation, "Read Delta Sharing Tables". https://docs.delta.io/delta-sharing/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data