Skip to content

Software companies

Marketing automation vendors: subscriber data vs your own platform records

By SourceX Editorial · Updated

Short answer

Marketing automation platform data ownership splits two ways. Subscriber lists, contact properties, campaign content and engagement events belong to customers and are generally off limits for licensing. The vendor's own deliverability cases, abuse desk decisions, support history and engineering records are the licensable core once customer and recipient details are removed.

Key takeaways

  • Subscriber lists, message content and engagement events are customer data you process, not records you own.
  • Deliverability, abuse desk and compliance case histories capture expert judgment and are company records.
  • Recipient addresses hide inside your own tickets, bounce logs and complaint forwards, so preparation carries most of the effort.
  • Customer AI training switches and stricter API terms show where the market has settled: customers control their data.
  • Aggregated engagement benchmarks are the gray zone and need clear contract language and counsel's sign-off.

Which records belong to your customers?#

The records that belong to your customers are everything they and their subscribers put into the platform: subscriber lists, contact properties, segments, campaign and automation content, form submissions and engagement events such as opens and clicks. You hold them as a processor, and your DPA limits what you may do with them.

The largest platforms now say so in writing. HubSpot's updated developer terms state that customer data belongs to the customer, not to HubSpot or developers, and restrict using data accessed through its APIs to train AI models, with a carve-out for legitimate single-customer use. HubSpot also gives customers a switch that controls whether it uses their account data to train its own AI models.

Expect your customers and their legal teams to hold you to the same standard. A martech vendor that treats subscriber data as raw material for its own projects invites contract disputes and churn, whatever its terms technically allow.

Subscriber data vs platform records: the dividing line#

The dividing line runs between data about your customers' audiences and records about how your company runs the platform. The first is customer data; the second is your operating history, even though customer details are woven through it.

Use the table below as a first sort, then confirm each row against your subscription agreement and DPA.

Two rows deserve a second look. Engagement events feel like platform telemetry because your systems generate them, but they describe how specific people responded to a customer's message. Support transcripts feel like customer content, yet they record your team's diagnosis and become usable company records once customer details are removed.

Subscriber data vs platform records: the dividing line
RecordTypical ownerLicensing outlook
Subscriber lists and contact propertiesCustomerExcluded
Email, SMS and landing page content customers writeCustomerExcluded
Opens, clicks, unsubscribes and form fillsCustomer, about end recipientsExcluded; aggregates only with clear contract terms
Deliverability cases and remediation notesVendor, with customer identifiers insideCandidate after removing customer and recipient details
Abuse desk and compliance decisionsVendorCandidate with careful preparation
Support tickets and chat transcriptsVendor, with customer content insideCandidate after preparation
Engineering issues, code reviews and postmortemsVendorStrong candidate
Product decisions and roadmap discussionsVendorCandidate

Why deliverability and compliance records stand out#

Deliverability and compliance records stand out because they capture expert decisions that are hard to find anywhere else. A deliverability engineer diagnosing why a sender's mail moved to spam folders, planning an IP warm-up or working through a blocklist removal leaves a trail of reasoning, actions and outcomes.

Abuse desk cases follow the same shape. An analyst reviews a spike in complaints, inspects how a list was collected, decides whether to suspend sending, and documents the customer's remediation. Support escalations about authentication records, bounce classification and API limits show the same cause, action and result pattern.

In SourceX Enterprise Data Value Framework terms, these records score on domain expertise and human-generated signal, and they are hard to reproduce from public sources. Preparation cost is the offsetting driver, because nearly every case mentions a customer domain or a recipient address.

Where recipient data hides inside your own records#

Recipient data hides inside your own records wherever a customer or engineer pasted evidence into a case. Removing it is the main preparation task for a martech vendor, and it should be planned before any scoping conversation starts.

A typical approach replaces addresses and domains with consistent placeholders, strips message bodies and attachments, and keeps the analyst's reasoning and the outcome. Consistent placeholders preserve the thread of a case without exposing who was involved.

  • Email headers pasted into tickets, which carry recipient and sender addresses.
  • Bounce and complaint logs attached to deliverability cases.
  • Complaint forwards from mailbox providers that quote the original message.
  • List exports customers attached when asking for import help.
  • Screenshots of contact records, segments and campaign reports.
  • Customer sending domains and account names that identify the customer.

Contract terms that decide the gray zone#

The contract terms that decide the gray zone are the usage data clause, the AI features clause and the DPA's purpose limits. Together they say whether you may keep aggregated statistics, whether customer data may improve your models, and whether anything derived from customer data can leave the company.

Aggregated engagement benchmarks are the hardest case. If you rely on de-identification, check the standard that applies: under California law as amended by the CPRA, deidentified information must not reasonably be linkable to a consumer, and the business must take reasonable measures against association, publicly commit not to reidentify it, and contractually bind recipients to the same terms.

If your terms are silent, adding new rights through a quiet terms update is risky, especially with enterprise customers on negotiated agreements. A signed amendment or an opt-in addendum is slower but holds up when a customer's counsel reads it.

Illustrative example: an email platform for independent retailers#

Illustrative: a fictional email and SMS platform serving independent retailers runs Zendesk for support, Jira and GitHub for engineering, an internal deliverability console, and a Slack channel where the abuse desk discusses suspensions. Its CEO wants to know what could be licensed to an AI developer building email operations agents.

The review excludes subscriber lists, message content and engagement events entirely, and sets aside engagement benchmarks because the usage data clause covers only service improvement. The candidate package is deliverability cases, abuse desk decisions, support escalations and engineering postmortems, with every address, domain and account name replaced by placeholders and message bodies removed.

The CEO approves that scope, and the company adds a plain statement to its trust page that subscriber data is never licensed. The statement answers the question customers ask most and matches what the contracts already say.

How SourceX scopes a martech vendor#

SourceX scopes a martech vendor by excluding subscriber data and message content in the Rights step of the SourceX five-step transaction, then assessing the vendor's own operating records for fit. The fit check asks only for metadata, such as which systems hold deliverability and support history and how many years remain accessible.

If a package proceeds, Preparation replaces customer and recipient identifiers, the vendor approves the result, and the SourceX Evidence Packet documents provenance, licensing rights, permitted use, the privacy record and release authorization.

Frequently asked questions

Can we license aggregated engagement benchmarks?

Only if your contracts clearly permit that use and the aggregates cannot be linked back to a customer or recipient. Service-improvement language usually does not stretch to licensing outside the company. Review the usage data clause, the DPA and any privacy commitments on your website with counsel, and expect some customers to object even where the contract allows it.

Do customer AI training opt-outs affect licensing our own records?

Usually not for records that are genuinely yours, such as engineering issues and internal decisions. They can matter for support tickets and cases that contain customer content, because a customer who opted out may expect that content to stay out of any training. Many vendors honor opt-outs across all customer-derived content to keep the position simple.

Are email templates we designed ourselves our records?

Stock templates your designers created are vendor property. Once a customer edits a template and fills it with its own copy, brand and offers, that version is customer content. Keep the stock library and its change history separate from customer copies so the line is easy to show.

What about agencies that resell our platform to their clients?

Reseller and agency arrangements add a layer: the agency may be your customer while its clients own the subscriber data. Read the reseller agreement and the end-client terms together, and treat records from agency accounts with extra care during preparation, because they often name several businesses at once.

Do anti-spam and consent laws affect what we can license?

They shape the records more than the license. Laws such as CAN-SPAM, and consent rules for text messaging, govern how your customers send, and your compliance cases document how those rules were applied. The cases still contain personal data about recipients, so the privacy laws that may apply are assessed deal by deal with counsel before any scope is approved.

Sources

  • HubSpot's developer changelog says its updated Developer Terms restrict using data accessed through HubSpot APIs to train, fine-tune or improve AI or machine learning models, with a carve-out for legitimate single-customer use cases. The same post says the terms now state that customer data belongs to the customer, not to HubSpot or developers. Source
  • HubSpot's knowledge base says customers can use an AI model training switch to control whether HubSpot uses their account's customer data to train its AI models. Source
  • Under Cal. Civ. Code § 1798.140(m), as amended by the CPRA, information is 'deidentified' only if it cannot reasonably be used to infer information about, or otherwise be linked to, a particular consumer. The business holding it must also (1) take reasonable measures to ensure it cannot be associated with a consumer or household, (2) publicly commit to keep and use it only in deidentified form and not try to reidentify it, and (3) contractually obligate any recipients to comply with all of these provisions. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify