Customer support ticket datasets for AI training
A customer support ticket dataset is a de-identified export of real support cases: the customer's messages, agent replies, internal notes, escalations, status changes and the final resolution, often linked to CSAT and QA scores. SourceX sources these histories from established companies running Zendesk, Intercom, Salesforce Service Cloud, ServiceNow or Freshdesk, typically covering two to ten years of tickets, and licenses them with an agreed scope and permitted use.
Dataset manifest
Sourced to your spec- What it is
- Resolved support cases with full threads, internal notes and outcomes
- Typical systems
- Zendesk, Intercom, Salesforce Service Cloud, ServiceNow, Freshdesk
- Typical history
- Typically 2–10 years of tickets; varies by partner
- Modality
- Text threads plus structured ticket metadata; attachments on request
- Delivery formats
- Agreed per order; JSONL or Parquet threads are common
- Preparation
- Names, contact details, order and account numbers de-identified before delivery
- Licensing
- Non-exclusive, or exclusive for an agreed snapshot where a program requires it
- Availability
- Sourced to your spec from partner companies; not guaranteed
What a delivery contains
Fields vary by source system and are fixed per order. A typical delivery includes:
| Field | Type | What it holds |
|---|---|---|
| ticket_id | string | Pseudonymous, stable case identifier, consistent across all tables in the delivery. |
| created_at | timestamp | When the case was opened, in UTC. |
| channel | enum | How the case arrived, such as email, chat, messaging, web form or phone with transcript. |
| product_area | string | The partner's own category, tag or queue for the issue. |
| priority | enum | Priority as set by the support team; later changes appear in the event log. |
| requester | object | Pseudonymous customer ID plus non-identifying attributes such as plan tier or region. |
| events | array | The ordered thread — customer messages, agent replies, internal notes, macros, status changes and escalations. |
| events[].author_role | enum | Who acted — customer, tier-1 agent, specialist, bot or system — without personal identity. |
| escalation_path | array | Teams or tiers the case moved through, in order. |
| resolution | object | Disposition code, reopen count, first-response and resolution times. |
| csat | object | Satisfaction score and free-text comment, where the customer was surveyed. |
| qa | object | QA scorecard by rubric dimension, with the rubric version used. |
| kb_links | array | Help-center or internal articles referenced while resolving the case. |
Example record
{
"ticket_id": "t_7f3c9a",
"created_at": "2023-03-14T09:12:44Z",
"channel": "email",
"product_area": "billing",
"priority": "normal",
"requester": { "id": "cust_51ad", "plan": "business" },
"events": [
{ "t": "2023-03-14T09:12:44Z", "type": "customer_message",
"text": "I was charged twice for invoice [INVOICE_ID] on my card ending [REDACTED]." },
{ "t": "2023-03-14T09:40:02Z", "type": "internal_note", "author_role": "agent_t1",
"text": "Duplicate capture visible in billing console. Refund over limit, needs T2." },
{ "t": "2023-03-14T09:41:10Z", "type": "macro_applied", "macro": "billing_duplicate_ack" },
{ "t": "2023-03-14T11:05:37Z", "type": "escalation", "from": "tier_1", "to": "billing_t2" },
{ "t": "2023-03-14T13:22:18Z", "type": "agent_reply", "author_role": "agent_t2",
"text": "We've refunded the duplicate charge. It will show in 3–5 business days." },
{ "t": "2023-03-14T13:22:30Z", "type": "status_change", "to": "solved" }
],
"escalation_path": ["tier_1", "billing_t2"],
"resolution": { "code": "refund_issued", "reopened": 0,
"first_response_min": 28, "resolution_min": 250 },
"csat": { "score": 5, "comment": "Fast fix, thanks" },
"qa": { "rubric": "2023.1", "accuracy": 4, "tone": 5, "process": 3 },
"kb_links": ["kb_duplicate_charges"]
}Synthetic record for illustration. Field names, structure and format are agreed per order.
What AI teams use it for
Fine-tune support agents on resolved cases
Real threads show how experienced agents diagnose a problem, which questions they ask first, and what a resolved case looks like under the company's actual policies.
Build held-out evaluation sets
Cases with known resolutions and QA scores become graded, private eval items that have never appeared on the public internet, so scores are not inflated by contamination.
Train routing and escalation models
Escalation events and tier changes label which cases needed a specialist, how long they waited, and where they ended up.
Ground agent-assist and retrieval
Tickets linked to macros and knowledge-base articles show which reference content actually resolved which kind of problem.
Model satisfaction and quality
CSAT and QA scorecards let you learn what separates an excellent resolution from an adequate one, beyond whether the ticket was closed.
Use-case guides: Customer support agents, Private evaluation sets, Voice agents and speech models
What makes this data valuable
Outcome labels
Disposition codes, reopen counts and resolution times tell you how each case ended.
Human judgment
QA scorecards and CSAT add graded feedback on real work rather than synthetic preference pairs.
Internal notes
Notes and escalation reasons record the reasoning that never reaches the customer.
Continuous history
Several years of tickets capture product changes, policy updates and seasonal spikes.
Linked knowledge
Macro and article references tie each answer to the material behind it.
Channel mix
Email, chat and messaging threads behave differently; several channels widen coverage.
Why real resolution histories are hard to substitute
Public support corpora are mostly short question-and-answer pairs from forums and FAQ pages. They show the customer's side and a polished answer, but not the work in between. Synthetic tickets can fill volume, yet they inherit the generator's assumptions about how support works, and they carry no ground truth about what actually resolved a case.
What an operating support team records is different. The internal note that says a refund is over the tier-1 limit, the macro an agent chose and then edited, the escalation to billing, the second reply that fixed it, the QA reviewer's score on process adherence: that sequence is the behavior an AI support agent has to learn, under constraints that are specific to each company. It only exists inside help-desk systems, and only in companies that have run support at volume for years.
That is also why these histories make strong evaluation data. A held-out set of real cases with known outcomes and QA scores measures whether an agent would have resolved the problem the way a strong human did, and none of it has been published online for a model to memorize.
How SourceX scopes a support dataset
A request starts with your spec: domain and product type, channels, languages, approximate volume, years of history, required labels, format and de-identification requirements. SourceX matches that spec against partner companies and builds a private manifest for each candidate dataset — source systems, years covered, approximate volume, what each record includes, how personal data will be handled, the status of licensing rights and availability.
You review the manifest and a sample before anything is agreed. Common adjustments at this stage are narrowing to specific queues or product areas, excluding channels with weak de-identification results, and requiring QA or CSAT coverage above a threshold. Scope, permitted use, exclusivity and price are then fixed in the license, and data moves only after the partner approves delivery.
What to check before licensing
- Confirm the company owns the support data. When a BPO or contact center handles support for a client, the client usually owns the tickets and must authorize licensing.
- Check that the partner's customer terms and privacy notices allow the agreed use, and how that review was done.
- Review de-identification on a sample, including free text, email signatures, screenshots and attachments.
- Measure label coverage on the sample — the share of cases with QA scores, CSAT or disposition codes.
- Compare language, channel and product-area mix against the distribution your model needs.
- Ask for macro IDs so you can deduplicate boilerplate replies that would otherwise dominate training.
- Agree permitted use, exclusivity, duration, derivative rights and deletion obligations in the license.
How licensing works through SourceX
- 1
Define
Send the domain, modality, volume, format, timeline and permitted use you need.
- 2
Source
SourceX identifies businesses that hold matching data and are open to licensing it.
- 3
Qualify
Fit, rights and quality are checked, and you review samples before committing.
- 4
License
Scope, permitted use, exclusivity, price and obligations are agreed in writing.
- 5
Deliver
Approved data is prepared, de-identified where required and transferred securely.
Questions buyers ask
Can I license real customer support tickets for AI training?
Yes, when the company that owns them agrees to license them. SourceX sources support histories from established companies, reviews licensing rights, de-identifies personal data and agrees scope and permitted use with both sides before anything is delivered. Availability depends on which partners hold matching data, so describe your target domain, channels, languages and volume in a data request.
How is personal data removed from support tickets?
Personal data is de-identified before delivery. That typically covers names, email addresses, phone numbers, order and account numbers and payment details, in structured fields and in free text, and attachments containing personal data are removed. Whether identifiers are redacted or replaced with consistent pseudonyms is agreed during scoping, so include your requirements in the request.
Can support data from outsourced (BPO) teams be licensed?
Only with the data owner's authorization. When a contact center or BPO runs support for a client, the client usually owns the tickets and the customer relationships. SourceX verifies ownership and authorization before such data is offered, which can add time to clearing a BPO-sourced dataset.
How many years of ticket history are typical?
Partners typically hold two to ten years of tickets. Longer histories capture product changes and policy updates, while recent years tend to have the most consistent tagging and QA coverage. You can set a historical timeframe in your request.
Can I review a sample before licensing?
Yes. Sample review is part of qualification: you check fit, rights and quality on a sample before scope, price and permitted use are agreed in the license. If the sample shows gaps, such as thin label coverage or residual personal data, the scope can be narrowed or the candidate dataset dropped.
Can I get exclusive access to a support dataset?
Sometimes. Some programs offer exclusivity for an agreed dataset snapshot or permitted use for a defined period. Exclusivity is negotiated in the license and usually affects price. Under a non-exclusive license, the partner remains free to license the same history to other buyers.
How much does a customer support dataset cost?
There is no public price list. Price depends on volume, years of history, label coverage such as QA, CSAT and outcomes, how specialized the domain is, the preparation work required and any exclusivity. The request form asks for a budget range so the scope can be matched to it.
Related datasets
- Contact center call recordings and transcripts
Recorded service and support calls with diarized transcripts, dispositions and QA scores
- Human feedback and QA-scored work
Work items with scores, verdicts and corrections from the people who reviewed them
- SOPs, playbooks and internal knowledge bases
Written procedures with page history, ownership and links to execution records
- IT service management and incident histories
Incidents, problems, changes and requests with work notes, CI links and outcomes
- Enterprise workflow and task execution histories
Linked task trajectories from request to outcome, across every tool the work touched
Evaluating this data for procurement?
Diligence packets are prepared per dataset. Rights, privacy processing and quality differ between datasets.
Request dataset diligenceTell us what your models need
Send your spec — domain, volume, format, timeline and permitted use — and SourceX will match it against partner data and come back with what can be licensed.
Updated 3 October 2026. Own data like this? See how companies license it to AI developers.