Skip to content

Software companies

Cybersecurity software vendors: what security data you can license

By SourceX Editorial · Updated

Short answer

A cybersecurity software vendor can most readily license its own records: detection engineering history, analyst triage decisions, playbooks, support cases and code reviews. Customer telemetry from endpoints, networks and logs is governed by customer contracts and privacy law, and threat intelligence sits in between, depending on how it was derived and whether third-party feeds restrict redistribution.

Key takeaways

  • Sort security data into customer telemetry, threat intelligence and vendor-owned records before any licensing discussion.
  • Customer telemetry usually belongs to the customer and often contains personal data such as usernames, hostnames and IP addresses.
  • Threat intelligence from third-party feeds commonly restricts redistribution and should be excluded.
  • Security records need extra preparation: secrets, network details and exploit code must be found and removed.
  • Licensing proprietary detection content without confidentiality terms can weaken its trade secret protection.

What security data does a vendor actually hold?#

A security software vendor holds three kinds of data that look alike in a data lake but differ sharply in ownership. Customer telemetry comes from the customer's endpoints, networks, identities and cloud accounts. Threat intelligence describes attackers and their tools. Vendor-owned records document how the vendor's own people built detections, triaged alerts, answered customers and shipped code.

Endpoint detection and response, managed detection, email security, SIEM and vulnerability management products all mix the three. The licensing question is never whether the vendor holds the data, but which class each record belongs to.

Customer telemetry, threat intelligence and vendor records compared#

The comparison below sorts the record types a security vendor typically holds by who controls them and how a licensing review usually treats them. Use it to brief leadership before anyone promises a dataset.

Customer telemetry, threat intelligence and vendor records compared
Data classExamplesWho usually controls itLicensing view
Customer telemetryProcess events, network flows, authentication logs, email metadataThe customer, under its contract with the vendorGenerally excluded; customer contracts and privacy law govern
Derived detections and alertsAlerts raised in customer environments, with contextMixed; depends on contract termsOnly with clear rights and heavy preparation
Threat intelligence from your own researchIndicators, malware analysis, attacker behavior write-upsThe vendor, if not derived from identifiable customer dataCandidate, subject to confidentiality and sharing rules
Third-party threat feedsCommercial or community indicator feedsThe feed provider or communityUsually excluded; redistribution limits apply
Detection engineering recordsRule histories, tuning decisions, false positive analysisThe vendorStrong candidate
Analyst and support recordsTriage notes, investigation playbooks, support casesThe vendor, containing customer detailsCandidate after customer details are removed
Engineering recordsIssues, code reviews, incident postmortemsThe vendorStrong candidate; scan for secrets first

Who owns EDR and product telemetry?#

EDR and product telemetry is usually customer data under the vendor's agreement, collected so the vendor can protect the customer. Many agreements give the vendor a license to use telemetry to provide and improve the service, and some add a right to derive threat intelligence. Few contemplated licensing telemetry to outside AI developers, and stretching an improvement clause that far invites a dispute.

Telemetry also carries personal data. Usernames, hostnames, file paths, email addresses and IP addresses can identify people, and under the GDPR, pseudonymized data that can be re-linked using additional information is still treated as information about an identifiable person. US state privacy laws may apply too, depending on whose data it is.

Many vendors publish responsible AI statements explaining how they train their own detection models on customer telemetry. Those statements describe the vendor's own use; they do not establish a right to give the same telemetry to a third party.

Threat intelligence: derived, shared and licensed in#

Threat intelligence sits between the other two classes because its origin varies. Indicators your researchers found by analyzing public malware are much easier to license than indicators extracted from a specific customer's incident, which may reveal that the customer was breached.

Shared intelligence carries its own rules. Information received through industry sharing groups often comes with Traffic Light Protocol markings or membership terms that limit onward sharing, and commercial feeds typically prohibit redistribution. Tag every indicator with its source so those limits travel with the data.

Vendor-owned records AI developers ask about#

Vendor-owned security records capture expert judgment, which is what model developers find scarce. The most useful ones link a signal to a decision and an outcome.

  • Detection engineering histories: the rule, why it was written, how it was tuned and what false positives taught the team.
  • Alert triage decisions with outcomes, from the vendor's own analysts in a managed service.
  • Investigation and response playbooks, including revisions made after real incidents.
  • Support cases where customers and engineers worked through deployment, tuning and integration problems.
  • Code reviews, issue histories and postmortems from the product's own engineering team.
  • Vulnerability disclosure handling records, with researcher and customer details removed.

Preparation risks unique to security data#

Security records need more preparation than most because they are full of things attackers want. Logs and tickets routinely contain credentials, API keys, internal hostnames and network layouts, and engineering records may contain working exploit code or detection logic that reveals how to evade the product.

Secret scanners help but need care. TruffleHog, an open-source scanner, covers sources including Git, chats, wikis and logs, and for secrets it can classify it can log in to confirm whether a credential is live, which means sending authentication requests. Run that verification only where you are authorized to.

Trade secret protection is the other risk. Under the Defend Trade Secrets Act, information qualifies as a trade secret only if its owner takes reasonable measures to keep it secret, so any license of detection content should carry confidentiality, access and use restrictions.

Preparation risks unique to security data
RiskWhere it hidesPreparation step
Secrets and credentialsLogs, tickets, chat, code and configuration filesScan with secret detection tools; remove findings and rotate live credentials
Customer network detailsHostnames, IP ranges and asset names in alerts and notesReplace with consistent placeholders
Breach facts about customersIncident tickets and investigation notesRemove customer identity and identifying incident details
Exploit code and evasion detailResearch repositories and detection rulesExclude, or limit by agreement with the licensee
Third-party intelligenceIndicators copied from feeds into notesTag by source and exclude restricted items

Illustrative: an MDR platform vendor separates its packages#

Illustrative: a fictional managed detection and response platform vendor serves mid-sized manufacturers and distributors. An AI developer building security operations agents asks about licensing triage and investigation records.

The vendor excludes customer telemetry entirely and drops indicators that came from paid feeds. The package that proceeds holds de-identified analyst triage notes with outcomes, the history of detection rules with tuning rationale, investigation playbooks and support cases. Preparation replaces customer hostnames and IP addresses with placeholders, removes every customer name and runs secret scanning across notes and code.

The license restricts use to training and evaluating models, prohibits publishing detection logic and requires the developer to protect the material as confidential.

How SourceX handles security vendor records#

SourceX begins with a metadata-only fit check; nothing is shared during the initial assessment. In the SourceX five-step transaction, the Rights step sorts records into telemetry, threat intelligence and vendor-owned classes, Preparation removes secrets and customer details, and the vendor approves the release.

Large security datasets stay in the vendor's own storage or ship on encrypted drives, and SourceX never hosts multi-TB datasets. Each package carries a SourceX Evidence Packet covering provenance, licensing rights, permitted use, the privacy record and release authorization.

Frequently asked questions

Can we use customer telemetry to train our own detection models?

Often, if your customer agreement grants a license to use telemetry to provide and improve the service, which is how many security products work. That is a different question from licensing telemetry to an outside developer, which most agreements do not clearly allow.

Are IP addresses and hostnames personal data?

They can be. Under the GDPR and some US state privacy laws, identifiers that can be linked to a person may count as personal data, and pseudonymized identifiers remain personal data if they can be re-linked. Treat them as personal data during preparation unless counsel concludes otherwise.

Does licensing detection content help attackers?

It can if the content is published or leaks, which is why licenses for detection content restrict publication, limit access and require confidentiality. Some vendors also hold back rules for current, high-value detections and license older or retired content instead.

What about records from MSSP partners who resell our product?

The partner agreement decides it. An MSSP may own its analysts' notes and its customer relationships even when the work happened inside your platform. Review the partner agreement before including any partner-generated records.

Does our SOC 2 report cover a licensed dataset?

Not by itself. A SOC 2 report is an attestation by a CPA firm about controls over a defined system, not a certification of a dataset. A licensee may still ask about your controls, so describe separately how the package was prepared and delivered.

Sources

  • GDPR Recital 26 states that personal data which have undergone pseudonymisation, and could be attributed to a natural person using additional information, should be considered information on an identifiable natural person. Source
  • TruffleHog is an open-source secret scanner that scans sources including Git, chats, wikis and logs, and can log in to confirm whether a classified secret is live. Source
  • Under 18 U.S.C. 1839(3), information qualifies as a trade secret only if the owner has taken reasonable measures to keep it secret and it derives independent economic value from not being generally known. Source
  • SOC 2 examinations use the AICPA Trust Services Criteria and result in an attestation report by a CPA firm, not a certification. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify