Skip to content

Software companies

Zero data retention promises: what B2B software vendors must be able to prove

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

To prove a zero data retention promise, a B2B software vendor needs five kinds of evidence: the AI provider's terms or zero-retention approval, configuration records showing it is switched on, logs and settings proving your own systems do not keep prompts, a current sub-processor list and an independent audit report. A promise you cannot evidence becomes a contract risk.

Key takeaways

  • Zero data retention and no training on customer data are different promises, and each needs its own evidence.
  • Your own logging, tracing and error tools break zero-retention promises more often than the AI provider does.
  • A promise is only as strong as each sub-processor's commitment behind it, so keep that chain documented.
  • Changing an AI data promise later needs clear notice; FTC staff have warned that quiet, retroactive terms changes may be unfair or deceptive.
  • A no-training promise to customers does not stop a vendor licensing its own records, such as code and internal engineering history.

What a zero data retention promise actually covers#

A zero data retention promise says that customer inputs and outputs sent to an AI feature are processed and then not stored, by the model provider or by you. It is narrower and stricter than a promise not to train on customer data, which allows storage but forbids using the data to improve models.

Vendors scope these promises in different ways, so read your own wording precisely. Notion, for example, states that its LLM providers use zero data retention by default for Enterprise plan workspaces, while for other plans they retain customer data for 30 days or fewer before deletion. A promise that depends on plan, region or feature needs evidence for each variant.

Write down what your promise covers in plain terms: which features, which data, which providers, what exceptions and from what date. That document is the index for every piece of evidence that follows.

The evidence list: what to keep and where it comes from#

The evidence for a zero data retention promise comes from five main sources, plus the contracts that tie them together. Keep each item dated and owned by a named person, because security reviewers ask when it was last checked as often as what it says.

The evidence list: what to keep and where it comes from
EvidenceWhat it showsWhere it comes fromRefresh when
Provider terms and zero-retention approvalThe model provider does not store inputs or outputsProvider agreement, data addendum, written approvalProvider terms change or you add a provider
Configuration recordsThe setting is enabled on the accounts and projects in useProvider console exports, infrastructure-as-code, change ticketsAny change to keys, projects or endpoints
Logging and observability settingsYour systems do not persist prompts or completionsAPM, tracing, error tracking and log pipeline configurationNew tools, new features or new log fields
Sub-processor listEvery party that touches the data and its commitmentDPA schedule and vendor due diligence filesBefore any new sub-processor goes live
Independent audit reportControls were designed and operated as describedSOC 2 report and bridge lettersEach reporting period
Customer contract clauseWhat you promised, with exceptionsMaster agreement, AI addendum, DPAEach new template or negotiated change

Where zero-retention promises usually break#

Zero-retention promises usually break inside the vendor's own stack, not at the model provider. Engineers add tracing to debug a slow AI feature, an error tracker captures the request body, or an analytics event records the prompt length and the prompt with it.

Derived data is the second gap. Embeddings, caches and evaluation sets built from production traffic can all hold customer content after the request ends. Notion's AI security page, for instance, gives embeddings their own rule, saying they are removed from its vector databases within 60 days of a page or workspace being deleted, which shows how a retention promise has to name each derived store.

  • Request and response bodies in application logs or API gateway logs.
  • Span attributes in distributed tracing and payloads in error tracking tools.
  • Prompts pasted into support tickets, Slack threads or Jira issues during debugging.
  • Vector stores, semantic caches and retry queues.
  • Evaluation and regression datasets sampled from production.
  • Fallback providers used during outages that sit outside the approved list.
  • Provider exceptions for abuse monitoring or legal holds that the customer promise does not mention.

Sub-processors and the chain of promises#

Sub-processor commitments are the links in the chain behind your promise, and customers increasingly ask to see each link. GitLab's Duo documentation, for example, states that GitLab does not train generative AI models and that all of its AI model sub-processors are restricted from using model input and output to train models. Atlassian's AI trust page states that its third-party-hosted LLM providers do not use customer inputs and outputs to improve their services.

Statements like these show the pattern customers expect: a public commitment backed by contract terms with each sub-processor. Keep the contract excerpt next to each entry on your sub-processor list, and follow the notice process in your DPA before adding a new AI provider.

An independent audit report supports the chain but does not replace it. SOC 2 examinations use the AICPA Trust Services Criteria, and a Type 2 report tests whether controls operated effectively over a period. Whether that covers your AI retention controls depends on the report's scope, so check the system description before relying on it.

Changing the promise later#

Changing an AI data promise later is where vendors take the most reputational and legal risk. FTC staff took up the point in a post dated February 13, 2024, cautioning that loosening data practices, for instance to begin training AI on consumers' data, and disclosing the change only by slipping it retroactively into a terms of service or privacy policy update could amount to an unfair or deceptive practice. The post addresses consumer data, but business customers apply the same test in security reviews and renewals, and contract law may apply to any change to negotiated terms.

Customers also read the fine print. TechCrunch reported in May 2024 that Slack drew backlash when its privacy principles were found to allow customer data to train its machine-learning models unless an organization opted out, and Slack responded that it does not train its generative AI models on customer data. Version your AI terms, give notice before changes and keep a record of which customers accepted which version.

Illustrative: a security review finds prompts in a trace#

Illustrative: a fictional SaaS company sells work order software to commercial property maintenance teams and has added an AI feature that summarizes technician notes. Its contract template promises zero data retention for that feature. An enterprise prospect's security team asks for evidence.

Assembling the evidence, the platform team finds that a tracing library records the full prompt as a span attribute, retained under the observability vendor's default settings. They remove the attribute, purge stored traces, add a test that fails the build if prompt fields appear in telemetry and record the fix in a change ticket.

The final evidence pack contains the provider's zero-retention approval, a console export, the tracing configuration, the sub-processor list with contract excerpts and the current SOC 2 report. The prospect accepts it, and the team adds the telemetry test to its release checklist.

Can a vendor with a no-training promise still license data?#

A vendor with a zero-retention or no-training promise can often still license records it owns, such as source code, code reviews, Jira issues and internal engineering discussions, because those promises cover customer data, not the vendor's own work. Customer content stays excluded unless the contract clearly permits otherwise.

In the SourceX five-step transaction (Supply, Rights, Preparation, Approval, Delivery), zero-retention and no-training commitments are read during Rights as limits on scope, and Preparation strips personal and confidential details from what remains. The vendor's SourceX Evidence Packet then documents provenance, licensing rights, permitted use, the privacy record and release authorization, which lets the vendor answer a customer's question about its data with a document rather than an assurance.

Frequently asked questions

Is zero data retention the same as not training on customer data?

No. Zero data retention means inputs and outputs are not stored after processing. A no-training promise allows storage but forbids using the data to improve models. Many vendors make both promises, and each needs its own evidence and its own contract wording.

Does a SOC 2 report prove zero data retention?

Not by itself. A SOC 2 report describes controls within a defined scope and, in a Type 2 report, tests whether they operated over a period. It supports your promise only if the system description and tested controls actually cover AI features, logging and sub-processors.

What should a zero data retention contract clause include?

It should define the data covered, the features and providers in scope, any exceptions such as abuse monitoring or legal holds, sub-processor commitments, notice before changes, audit or evidence rights and what happens if the promise is breached. Vague clauses invite disputes.

How often should the evidence be refreshed?

Refresh it whenever something it describes changes: a new provider, a new AI feature, a new observability tool or new contract wording. Many vendors also review the whole pack on a fixed cycle tied to their audit period, so evidence is current when reviewers ask.

What if we discover our promise was not met?

Treat it as an incident. Stop the retention, purge stored copies, document scope and timing and take advice on contractual and legal notice duties. Customers tend to judge the response as much as the lapse, so be accurate and prompt.

Sources

  • Notion states that LLM providers use zero data retention by default for Enterprise plan workspaces, while for non-Enterprise workspaces LLM providers retain Customer Data for 30 days or fewer before deletion; embeddings stored in vector databases are deleted within 60 days after a page or workspace is deleted. Source
  • GitLab's Duo data usage documentation states that GitLab does not train generative AI models and that all GitLab AI model sub-processors are restricted from using model input and output to train models. Source
  • Atlassian's AI Trust page states that Atlassian does not share customer metadata or in-app data with its third-party-hosted LLM providers for them to train or improve their services, and that those LLM providers do not use customer inputs and outputs to improve their services. Source
  • SOC 2 examinations use the AICPA's 2017 Trust Services Criteria, and a Type 2 report also tests whether controls operated effectively over a specified period. Source
  • On February 13, 2024, FTC staff warned that a company adopting more permissive data practices, such as using consumers' data for AI training, and telling consumers only through a surreptitious, retroactive change to its terms of service or privacy policy may be engaging in unfair or deceptive practices. Source
  • TechCrunch reported on May 17, 2024 that Slack drew user backlash after its privacy principles were found to allow customer data to be used to train Slack's machine-learning models unless an organization opted out, and Slack responded that it does not use customer data to train its generative AI large language models. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify