Skip to content

Industry-specific operational data

SaaS support tickets linked to bug reports and fixes for AI

Quick answer

A support-ticket-to-bug-report dataset joins technical support cases from a help desk such as Zendesk, Salesforce Service Cloud or Intercom to the engineering issues they were escalated into, usually in Jira, Linear or GitHub Issues, along with the product version the customer ran, the issue's status history, any workaround and the release that shipped the fix. That join is what teaches a model when to escalate, which known bug a new case matches, and which answer applies to which version.

By SourceX Editorial · Updated

What the linked record must contain

The linked record is only useful if the case, the issue and the release timeline survive the export together. A ticket export alone gives you text and a resolution code; it does not tell you that 40 cases with different wording were the same defect, or that the right answer changed when version 7.3 shipped. Ask for the join key (the issue key stored on the ticket, often in a custom field or a Jira/Zendesk integration link table) and the timestamps on both sides.

Core fields to request, at minimum:

  • Case side: case ID, full thread with author role (customer, agent, tier-2, engineer), product area, plan tier, channel, reported version or build, environment (browser, OS, region, deployment type), priority, escalation timestamp and final case outcome.
  • Attachments: HAR files, console logs and stack traces, scrubbed before delivery (see below).
  • Issue side: issue key, type, component, affected versions, fix versions, status transitions with timestamps, duplicate and "relates to" links, and the linked pull request or commit where it exists.
  • Knowledge side: the workaround text and, ideally, the help-center or internal KB article version that was current when the case was answered.

For generic ticket corpora without engineering links, see the customer support ticket datasets page; for how escalation itself works operationally, see customer support escalation records.

Which models this data trains and evaluates

This data supports three distinct model families: escalation classifiers, duplicate-bug matchers and version-aware answer or agent systems. Each uses a different slice of the same join, so specify which one you are building before scoping.

  • Escalate or resolve: the label is whether tier-1 escalated, and whether engineering confirmed a defect versus a configuration or "works as designed" outcome. Cases escalated and then closed as user error are the most valuable negatives.
  • Duplicate detection and triage: many-to-one links between cases and a single issue give you positive pairs for free. Issue-tracker "duplicates" links add issue-to-issue pairs.
  • Version-aware RAG and tier-2 agents: the model must answer differently for a customer on 7.2 than on 7.4. Evaluating that requires the knowledge-base snapshot the answers came from, not only question-answer pairs; WixQA, an enterprise RAG benchmark built on support content, ships its questions alongside the knowledge-base snapshot used for grounding [1].

Public issue-to-PR benchmarks such as SWE-bench pair GitHub issues with the pull requests that resolved them [2], but they contain no customer-side case, no plan tier and no workaround. They are also exposed to contamination: as of October 2026, OpenAI says it no longer reports SWE-bench Verified scores because the benchmark is increasingly contaminated [3]. Private, licensed support-to-engineering histories are one way around both gaps. If your target is the code change rather than the case, the issue-to-fix pairs guide and stack-trace-to-fix data cover that side.

Link labels are noisy in predictable ways, so ask how each link was created before treating it as ground truth. The biggest question is whether a case-to-issue link was set by an agent who read both, by a tier-2 engineer during triage, or by automation such as a macro or keyword rule.

Common failure modes to test for in a sample:

  • Bulk linking during incidents. During an outage, every inbound case may be linked to one incident issue regardless of symptom. Flag links created within minutes of each other by one user.
  • Missing negatives. Agents link when they find a match and leave nothing when they do not, so an unlinked case is not proof of "no matching bug."
  • Stale fix versions. Fix versions are often edited after release; use the status history, not the current field, to reconstruct what was known at answer time.
  • Canned replies. Macros make thousands of agent turns near-identical, which inflates apparent dataset size and encourages memorization; deduplication is known to reduce memorization and train-test overlap [6]. See templates and boilerplate in business records.
  • Version field drift. "Reported version" is frequently free text ("latest," "the new UI"); ask what share is parseable to a real build.

Secrets and personal data in logs and attachments

Log attachments are the highest-risk part of this dataset, so secret scanning and PII scrubbing are mandatory before any training use. HAR files and debug logs routinely carry bearer tokens, session cookies, API keys, customer emails, tenant IDs and payload fragments from the customer's own users.

Secrets matter beyond confidentiality: research on unintended memorization shows language models can memorize and later expose rare sequences such as secrets that appear in training data [4]. Enterprise studies also find credentials leaking into document-sharing and code platforms often enough to need dedicated detection [5]. Ask suppliers which scanners and rule sets were run on attachments, whether tokens were revoked or only masked, and how replaced values stay consistent across a thread so that "customer A's tenant" remains one entity.

When working through SourceX, personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method used is recorded and a sample is checked; no de-identification method is perfect, so your own review still matters.

Rights: customer agreements decide what can be licensed

The supplier owning its help desk does not settle whether customer-submitted content can be used for AI training; its customer agreements and DPAs do. Enterprise SaaS contracts often restrict use of customer data to providing the service, and FTC staff warned in a January 2024 post that companies may be liable if they use customer data for undisclosed purposes such as training models, contrary to their commitments [7].

Ask for the clause-level basis: which terms of service and DPA versions applied to the cases in scope, whether enterprise accounts with negotiated terms were excluded, and whether log payloads belonging to the customer's end users were dropped. The customer contracts and DPAs guide walks through that review. This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Request template for a linked support-and-bug dataset

A precise request names the join, the time window, the label you need and the rights questions up front. Adapt the template below; documentation fields follow the Data Cards pattern of recording sources, collection and annotation methods and intended use [8].

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample specification
Use caseDuplicate-bug matching and escalate/resolve classifier for a B2B SaaS support agent
Case sourceHelp desk export (JSON per ticket), full threads with author role
Engineering sourceIssue tracker export with status history, affected and fix versions, duplicate links
JoinIssue key on case; link creation method and creator role per link
Version contextReported build parsed to release; release calendar; KB article versions by date
Window24 months, with at least two major releases
AttachmentsLogs and HAR files included only after secret and PII scanning; scanner report attached
ExclusionsAccounts whose contracts bar AI use; end-user payload content
Labels wantedEscalated (y/n), defect confirmed, duplicate-of, workaround applied, fix version
DocumentationData card: sources, link provenance, scrubbing method, known gaps

Worked example of one joined record, also invented:

{
  "case_id": "C-118204",
  "plan_tier": "enterprise",
  "product_area": "SSO / SAML",
  "reported_version": "7.2.4",
  "thread": [{"role": "customer", "text": "SAML login loops after IdP cert rotation"}],
  "attachments": [{"type": "har", "scrubbed": true, "secrets_found": 3}],
  "linked_issue": {"key": "AUTH-2291", "link_created_by": "tier2_agent",
                   "status_history": ["Open", "In Progress", "Resolved"],
                   "fix_version": "7.3.0"},
  "workaround": "Re-upload metadata XML instead of cert only",
  "outcome": "resolved_with_workaround"
}

How this differs from neighboring datasets

This page covers the case-to-issue-to-release join; neighboring datasets cover one end of it. Ticket-to-knowledge-article pairs are relevance labels for retrieval and are covered in support tickets linked to knowledge articles. Incident tickets, runbooks and postmortems for SRE agents are covered in IT operations agent evaluation and the ITSM ticket datasets page.

Code-side histories (commits, reviews, fixes) are owned by software engineering datasets and software bug fixing records. Buyers comparing software data more broadly can start from software industry buyers or the industry-specific operational data hub.

How SourceX handles requests for linked support and engineering data

SourceX sources operational datasets, including support histories and engineering records, from US companies on request and manages the licensing and ongoing purchases. Nothing is held in stock and a request does not guarantee a match; you describe the data, SourceX looks for US businesses that hold it, and every release is approved by the supplying company. Datasets are rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery, through private, access-controlled workflows after an executed agreement. Describe the join, versions and labels you need on the SourceX buyer request page.

Request support tickets linked to bug reports

If you need cases joined to issues, versions and fixes, send SourceX a description of the fields, window and use. SourceX assesses data and licensing permissions with the supplier, and nothing is contracted until a supplier agrees. Start a buyer request.

Sources

  1. arXiv, "WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation" (2025). https://arxiv.org/html/2505.08643v1
  2. arXiv (ar5iv), "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://ar5iv.labs.arxiv.org/html/2310.06770
  3. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  4. arXiv, "The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks" (2018). https://arxiv.org/abs/1802.08232v2
  5. arXiv, "Using AI/ML to Find and Remediate Enterprise Secrets in Code & Document Sharing Platforms" (2024). https://arxiv.org/html/2401.01754v1
  6. arXiv, "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499
  7. Federal Trade Commission, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  8. arXiv (Google Research, FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data