Retrieval, RAG and grounding data
Support tickets linked to knowledge articles: ready-made relevance labels
Quick answer
When a support agent attaches or cites a knowledge-base article while solving a ticket, the ticket becomes a real customer query and the article becomes a judged positive passage. These links are implicit relevance labels, already recorded in the helpdesk. To use them for retriever training or RAG evaluation, buy three parts together: ticket text with timestamps, the article ID and version at link time, and a pinned snapshot of the knowledge base as it stood then. Expect noisy positives and plan to validate a sample.
By SourceX Editorial · Updated
Why ticket-to-article links work as supervision
Ticket-to-article links pair real user wording with the document a trained agent judged useful, which is the query-passage pair a support retriever needs. Synthetic questions generated from articles tend to reuse the article's vocabulary. Real tickets carry the mismatch a support retriever actually faces: "can't log in after the update" against an article titled "Resetting SSO sessions after a client upgrade."
Published work points the same way. A domain-specific RAG system for a software help center trained its dense retriever on logs that linked real user queries to help articles, treating real interactions as supervision [2]. WixQA, a benchmark built on an enterprise support knowledge base, shows the other half of the requirement: end-to-end evaluation needs the question-answer pairs and the specific KB snapshot the answers came from, which is why the authors release that snapshot [1].
Linked tickets also sit between two record types you might otherwise buy separately. Raw customer support ticket datasets and licensed knowledge base articles are each useful alone; the join between them is what turns them into a retrieval set.
Where the links live in common helpdesk systems
The links are recorded in different places depending on the helpdesk, and the extraction method decides how much signal survives. Ask the supplier which of these mechanisms produced each link, because they have very different noise profiles.
- Salesforce Service Cloud: the
CaseArticlejunction object links a Case to a Knowledge article version. It is a structured link, but agents sometimes attach articles in bulk at close. - ServiceNow: knowledge usage on incidents and cases is typically recorded in task-to-knowledge relationship and knowledge-use tables, and article references also appear in work notes and comments held in
sys_journal_field. - Zendesk: links usually appear as article URLs inside public replies or internal notes, sometimes inserted by the Knowledge Capture app. The Ticket Audits API exposes each comment and field change as an event with author and values, which lets you reconstruct when a link was added and by whom [4].
- Jira Service Management and Freshdesk: linked Confluence pages or solution articles appear as issue links or in reply bodies; URL parsing is often the only extraction path.
URL-parsed links need normalization. Help center URLs change with locale prefixes, slugs and redirects, so map every URL to a stable article ID before you count it as a label.
What a usable delivery contains
A usable delivery contains the ticket, the link event and the exact article text the agent saw, joined by stable IDs. Missing any one of the three breaks either training or evaluation.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Table | Key fields | Why it matters |
|---|---|---|
tickets | ticket_id, created_at, channel, product_area, first_customer_message, subject, status, solved_at | The first customer message is the query; later turns leak the answer |
ticket_events | ticket_id, event_at, author_role (agent, customer, bot), visibility (public, internal), body | Shows whether the article was sent to the customer or only consulted |
ticket_article_links | ticket_id, article_id, article_version, linked_at, link_source (structured, URL-parsed, macro, bot), link_position | The label itself, with provenance for filtering |
outcomes | ticket_id, resolution_code, reopened, csat_score, escalated | Lets you weight positives that actually resolved the issue |
kb_snapshot | article_id, version, locale, title, body_html, valid_from, valid_to, status (published, archived) | The corpus your retriever searches, frozen at link time |
Ask for article_version explicitly. Knowledge articles are edited constantly, and a label pointing to the 2024 text of an article whose current version covers a different product release is a false positive in disguise. If the supplier cannot provide version history, at minimum request a snapshot dated close to the ticket window and a list of articles edited inside it. For the general pattern, see pinned corpus snapshots for reproducible RAG evaluation and versioning licensed datasets.
Turning links into qrels without fooling yourself
Convert links into graded qrels only after separating strong links from weak ones, because agent-attached links are noisy positives. Common failure modes include macros that append the same "contact us" article to every reply, bulk attachment at ticket close to satisfy a knowledge-reuse metric, links to a generic troubleshooting hub rather than the specific fix, and bots that insert suggestions the customer never opened.
Label noise matters more in evaluation than in training. An audit of widely used test sets estimated an average label error rate of at least 3.3% and showed that such errors can change model rankings [5]. For a support eval set, a practical approach is:
- Drop links whose
link_sourceis a macro or bot unless the agent edited the reply around them. - Weight links by outcome: solved without reopen and with a positive CSAT counts as relevance grade 2; solved with reopen counts as grade 1.
- Treat the absence of a link as unjudged, not irrelevant. Agents rarely attach every article that would have helped.
- Have domain reviewers judge a stratified sample by product area and link source, and report precision per stratum.
Step 3 is where most teams go wrong. Ticket-derived qrels are sparse, so metrics like recall@10 understate a retriever that surfaces a correct but unlinked article. Pool top results from several systems and judge the pool; the BEIR authors note that judgments gathered mainly from lexical systems can understate other retrievers [3]. For assessor design, see relevance assessment guidelines and graded scales, and for when model judges can fill gaps, LLM relevance labels vs human assessors.
Building the evaluation split
Build the evaluation split by time, not at random, so the test tickets come after the training tickets and are scored against the snapshot that existed when they were opened. A random split lets near-duplicate tickets from the same incident spike land on both sides and inflates results.
Package the eval set in a format your harness already reads. A BEIR-style layout of corpus.jsonl, queries.jsonl and a qrels TSV covers retrieval metrics [3]. For end-to-end RAG scoring, add the agent's final public reply as a reference answer and map each record to the question, contexts, answer and ground-truth structure used by tools such as Ragas [6]. Keep the agent reply out of the training side if you plan to use it as a generation reference.
Also hold out a slice of tickets with no linked article. Those are your coverage gaps: queries the knowledge base did not answer. They feed coverage-gap analysis and test whether your assistant declines or escalates instead of inventing an answer.
De-identification that keeps the retrieval signal
De-identify customer text in a way that removes people while keeping product names, error codes and version strings, because those tokens carry most of the retrieval signal. A generic named-entity scrubber will often redact "Error 0x80070005" or a plan name as an identifier and quietly destroy the queries.
Ask the supplier to replace names, emails, phone numbers, account and order numbers with typed placeholders such as [EMAIL] or [ACCOUNT_ID], keep vendor product terms on an allowlist, and document the method. Watch for identifiers in attachments, signatures and pasted log excerpts, which often contain hostnames, IP addresses and tenant IDs. Sensitive categories can hide in support text too, from health details in a benefits-platform ticket to payment disputes; see special category data in operational training datasets.
Rights checks specific to support data
Support tickets are customer communications, so the supplier's right to license them depends on its own customer terms and DPAs, not only on ownership of the helpdesk instance. The FTC has warned AI companies that using customer data for undisclosed purposes such as model training, contrary to their privacy commitments, can create liability under laws it enforces [8]. Ask whether the supplier's terms of service and privacy notice cover the use, and whether any enterprise customers negotiated contracts that exclude their tickets.
Knowledge articles raise a separate question. They are usually the supplier's own copyrighted content, but some embed third-party vendor documentation or partner content. Confirm that the KB snapshot license covers indexing and embedding, not only training, and read grounding license vs training license and customer contracts and DPAs for training use before you sign.
Buyer checklist for a ticket-to-article dataset
Use this checklist to scope a request before you talk to any supplier, so you describe the data rather than a vendor.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Helpdesk and KB platform: which systems produced tickets and articles, and which link mechanism (structured junction, URL in reply, macro, bot).
- Window and volume shape: ticket date range, share of tickets with at least one link, product areas and locales.
- Link provenance:
link_source,linked_atand author role for every link. - Version fidelity: article version at link time, or a snapshot plus edit log for the window.
- Outcome fields: resolution code, reopen flag, CSAT, escalation, and how reliable each is; see verifying outcome labels in operational records.
- Privacy method: typed placeholders, product-term allowlist, sample QA results.
- Rights basis: customer terms covering the use, KB content ownership, carve-outs.
- Documentation: a data card covering sources, collection, link extraction and known noise [7].
- Validation sample: a small reviewed slice with precision by link source before full delivery.
If you want the bug-tracker version of this join, where tickets link to engineering issues and fixes rather than articles, see SaaS support tickets linked to bug reports. For broader enterprise search corpora that mix tickets with wikis and drives, see multi-source enterprise search corpora.
How SourceX handles ticket-to-article requests
SourceX sources operational datasets from US companies on request, including support histories and documents, and manages the licensing process. Nothing is held in stock, and a request does not guarantee a match. Buyers describe the data they need, such as the fields in the checklist above, and SourceX looks for US businesses that hold it; every release is approved by the supplying company.
Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe a ticket-to-article dataset to SourceX at any stage of scoping. More context on the use case is on knowledge retrieval evaluation, customer support AI training data, connecting ticket histories with documentation and the retrieval and RAG data hub.
Source ticket-to-article linked data for support retrieval
SourceX sources support histories and documents from US companies on request and manages the license, with rights review and de-identification before private, access-controlled delivery. Matches depend on suppliers agreeing, and nothing is contracted until they do. Describe the linked ticket and KB data you need.
Frequently asked questions
Can I use ticket-article links without the KB snapshot?
You can train a retriever on query-article ID pairs, but you cannot evaluate reliably without the article text as it stood at link time. Without the snapshot, scores mix retrieval quality with content drift [1].
How many linked tickets do I need for an eval set?
There is no fixed number. Size it so each product area and locale you care about has enough judged queries to report per-segment metrics with useful confidence, and favor reviewed quality over raw volume.
Are bot-suggested articles useful labels?
They are weak labels at best, because they reflect the old retriever's choices and would bias your evaluation toward it. Keep them as a separate stratum, or use them only as hard negatives where the customer still needed an agent.
Sources
- arXiv, "WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation" (2025). https://arxiv.org/html/2505.08643v1
- arXiv, "Retrieval Augmented Generation for Domain-specific Question Answering" (2024). https://arxiv.org/pdf/2404.14760
- arXiv, "BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models" (2021). https://arxiv.org/pdf/2104.08663
- Zendesk Developer Docs, "Ticket Audit events reference". https://developer.zendesk.com/documentation/ticketing/reference-guides/ticket-audit-events-reference/
- arXiv (NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Ragas documentation, "Prepare your test dataset". https://docs.ragas.io/en/v0.1.21/getstarted/prepare_data.html
- arXiv (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.