Skip to content

Software companies

Customer data vs company data: what a SaaS company can actually license

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

The line between customer data and company data decides what a SaaS company can license. Customer content belongs to customers and is usually off limits; service-generated data depends on the usage and aggregated-data clauses; company-created records, such as code reviews, internal docs and support resolutions, are usually the company's to license after privacy preparation.

Key takeaways

  • Sort every record family into customer content, service-generated data or company-created records before discussing any license.
  • Customer content is governed by the customer agreement and DPA and is rarely licensable without new consent.
  • Service-generated data sits in the middle; the exact usage and aggregated-data wording decides how far it reaches.
  • Company-created records are the usual starting point, but they often quote customer content that must be removed.

What is the difference between customer data and company data?#

The difference between customer data and company data is who created the record and which contract governs it. Customer data is what customers put into or create in your product; company data is what your own team creates while building, selling and supporting that product. Most SaaS terms state the first point plainly: Slack's Supplemental Terms, for example, say the customer retains all ownership of its Customer Data.

A third class sits between them. Service-generated data, such as logs, usage events, performance metrics and model outputs, is produced by your system as customers use it. Contracts treat it inconsistently, which is why it causes the most disagreement inside SaaS companies.

Thinking in three classes rather than two prevents the most common error in early licensing conversations: assuming that anything stored in your infrastructure is company data because your company pays for the servers. Hosting a record and controlling it are separate questions, and the contracts answer the second one.

The three-class framework for SaaS records#

The three-class framework for SaaS records sorts each record family by source and states its default position. Your own agreements set the final answer, but most SaaS contracts follow this pattern closely enough to make it a reliable first cut.

Read the table by column, not by row. Each class has its own deciding document and its own decision-maker, which is why the classes are reviewed separately even when the records sit in the same database.

The three-class framework for SaaS records
AttributeCustomer contentService-generated dataCompany-created records
ExamplesUploaded files, messages, records customers create in the productLogs, usage events, feature telemetry, error traces, aggregate metricsCode, code reviews, Jira issues, internal docs, support resolutions, sales notes
Default licensabilityRarely, without new customer consentSometimes, in aggregated or derived formUsually, after rights and privacy review
Clause that decidesCustomer Data definition, license grant, DPA, confidentialityUsage data and aggregated-data clausesEmployee and contractor IP terms, internal policies
Main preparationNot applicable unless customers opt inRemove identifiers and test re-identification riskRemove customer content, personal details and secrets
Who decidesEach customer, then your signerYour counsel, reading your own termsYour CEO or authorized signer

Where do the gray zones appear?#

Gray zones appear wherever a company-created record quotes, attaches or summarizes customer content. These records are often the most valuable in the archive, because they show your team solving a customer's real problem from report to fix.

The usual resolution is to keep your team's work and remove the customer's material. That preserves the diagnosis, reasoning and outcome that AI developers care about while leaving customer content where the contract says it belongs.

Where do the gray zones appear?
RecordWhy it is grayUsual resolution
Support ticket threadsYour agent's work wrapped around the customer's messageKeep the resolution; remove or summarize customer content
Customer-reported bugs in JiraReproduction steps may paste customer recordsStrip pasted data; keep diagnosis and fix
Shared Slack channels with customersBoth parties wrote the messagesUsually exclude; internal channels are cleaner
Sales call recordingsThe prospect's voice and statementsExclude, or use only with documented consent
Implementation workbooksYour template filled with customer configurationKeep the template; drop customer entries
Outputs from in-product AI featuresGenerated from customer inputsTreat as customer-derived unless the contract says otherwise

When can service-generated data be licensed?#

Service-generated data can be licensed when the contract gives the vendor rights in it and the licensed form no longer reveals any customer's content or identity. Error traces and application logs often fail the second test, because they capture fragments of what users typed or uploaded.

Aggregated metrics, such as how long a workflow step takes across all customers, pass more easily but carry little training value on their own. Detailed event streams carry more value and more risk. Counsel should read the usage data clause for its stated purposes, check whether it allows disclosure to third parties, and confirm that de-identified outputs cannot be linked back to a customer.

Why buyers care which class a record belongs to#

Buyers care which class a record belongs to because the class determines whether the supplier can prove it had the right to license the record. AI developers increasingly ask for provenance and permitted-use documentation before accepting a dataset, and a package that mixes classes without labels is slow to clear their review.

Published standards point the same way. The Data and Trust Alliance's Data Provenance Standards organize dataset metadata into Source, Provenance and Use groups, which the specification describes as necessary for proper dataset selection for AI model training. A clean three-class sort makes that kind of documentation straightforward to produce.

How to sort your record families#

Sorting record families takes a working session with the CTO, head of support and counsel, using metadata only. No exports are needed to make the first cut, and the result can be tested by counsel before any data moves.

The output is a short map of systems, record families, classes and governing clauses. That map later becomes the backbone of a data inventory and the rights section of any license.

  • List each system that holds records: GitHub or GitLab, Jira or Linear, Zendesk or Intercom, Salesforce or HubSpot, Slack, Confluence or Notion, and the product database.
  • For each system, name the record families it holds, such as pull requests, incident reports or ticket threads.
  • Assign each family to one of the three classes, and mark gray-zone families separately.
  • Write down the clause or policy that governs each class, citing the version of the customer terms in force when the records were created.
  • Flag records from acquired products, contractors and resellers for separate review.
  • Choose the company-created families with the deepest linked history as the first scope.

Illustrative: a landscaping software company draws the line#

Illustrative: a fictional maker of routing and crew scheduling software for commercial landscaping companies runs the sort. Its product database holds customers' property lists, crew schedules and service histories, all classed as customer content and left out.

Route optimization logs and feature usage events are service-generated. The customer agreement allows use of aggregated, de-identified usage data to improve the service, so counsel concludes they can inform the company's own product work but should not be licensed in raw form.

Company-created records carry the package: years of GitHub pull requests and code reviews on the routing engine, Jira issues linked to Intercom conversations, and Confluence design documents. Property addresses pasted into bug reports are stripped, and the company proceeds with that scope.

How SourceX applies the framework#

SourceX applies the three-class sort in the Supply and Rights steps of the SourceX five-step transaction, before any preparation work begins. The SourceX Enterprise Data Value Framework then rates the company-created records qualitatively on drivers such as uniqueness, domain expertise, human-generated signal, recency, rights and AI utility, weighed against preparation cost and privacy burden.

Each package that proceeds is documented in a SourceX Evidence Packet: provenance, licensing rights, permitted use, the privacy record and release authorization. The supplier approves every step, and the company keeps ownership because the records are licensed, not sold.

Frequently asked questions

Is usage telemetry company data?

Often, but not always. Many SaaS agreements give the vendor rights in usage data and aggregated statistics, while others define customer data broadly enough to include events tied to a customer's users. Read the definitions and any aggregated-data clause; where personal information is involved, privacy laws may also limit use.

Do we own records our contractors created, such as code from an outsourced team?

Only if the contractor agreement assigned the rights to you. Most well-drafted agreements do, but older or informal arrangements may not. Check each contractor and agency agreement for an IP assignment before including their code, documentation or tickets in a license.

Can customer data become company data after a contract ends?

Generally no. Most agreements require return or deletion of customer data at termination, and confidentiality obligations often survive. Keeping it for a new purpose after the relationship ends is one of the riskier moves a vendor can make.

What about records from a product we acquired?

They follow the acquired company's contracts and the acquisition agreement. The seller's customer terms may be stricter or looser than yours, and some rights may not have transferred cleanly. Review them as a separate group before mixing them with your own records.

Does the classification change if records are de-identified?

No. De-identification lowers privacy risk but does not change who controls the record. A customer's support message is still customer content after names are removed. The class depends on origin and contract, while preparation decides how the record is cleaned.

Sources

  • The Data and Trust Alliance's Data Provenance Standards (version 1.0.0) define dataset metadata in three groups, Source, Provenance and Use, and say this metadata is needed to enable proper dataset selection for AI model training. Source
  • Slack's Supplemental Terms state that the Customer retains all ownership of its Customer Data. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify