Industry-specific operational data
Telecom network trouble ticket data for AI troubleshooting agents
Quick answer
A useful telecom trouble ticket dataset links each fault report to what the network said and what fixed it: service type, customer-described symptom, automated line, modem or signal test results, scripted actions taken, escalation tier, the dispatch or no-dispatch decision, and the closing cause and resolution codes. Public agent benchmarks simulate this workflow, but only real tickets carry real fault distributions. Because ticket fields can be customer proprietary network information (CPNI), licensing needs de-identification and legal review.
By SourceX Editorial · Updated
Why synthetic telecom benchmarks are not enough for training
Synthetic environments are good for testing tool use but cannot teach a model the real mix of faults, noisy notes and wrong first guesses that carriers see. The tau2-bench suite added a telecom domain where both the agent and a simulated customer act through tools, for example toggling device settings while the agent reads account and line state, and the whole environment is generated [1]. Its predecessor, tau-bench, set the pattern of scoring agents against databases and written policies with a simulated user [2].
That design measures whether an agent follows a procedure. It does not tell you how often a "no sync" complaint on a DSL line is really a premises wiring fault, how often a modem reboot clears a DOCSIS T3 timeout, or how many tickets reopen within 30 days. Those base rates exist only in operational ticket histories from Remedy, ServiceNow, Salesforce Service Cloud or in-house OSS ticketing systems.
Use synthetic suites for regression testing and real tickets for supervised fine-tuning, classifier training and realistic held-out evaluation. Real tickets also make better seeds for user simulators that behave like actual customers.
Fields a telecom fault ticket record should carry
The minimum viable record links a symptom to test evidence, an action sequence and a verified outcome. Ask suppliers which of these fields are populated, how they are coded and how often they are blank.
- Service context: service type (FTTH/GPON, DOCSIS cable, VDSL, fixed wireless, mobile, business Ethernet, SIP trunk), plan tier, CPE model and firmware.
- Symptom: symptom or trouble code, free-text customer description, channel (IVR, chat, agent, self-service app), and reported start time.
- Automated tests: line test results (loop length, attenuation, SNR margin, sync rate), cable modem stats (upstream and downstream power, codeword errors, T3/T4 timeouts), PON optical power, and the timestamp of each test relative to ticket open.
- Actions: scripted troubleshooting steps executed, remote resets or reprovisioning, and agent notes per step.
- Routing: escalation tier, transfers between Tier 1, Tier 2 and NOC queues, and links to parent outage or incident tickets.
- Dispatch: truck roll yes or no, technician type, dispatch outcome (fixed, no trouble found, customer not home) and any repeat dispatch.
- Closure: cause code, resolution code, fault location (network, drop, inside wire, CPE, customer education), close time and a repeat-ticket flag within a set window.
If the ticket system exposes data through TM Forum's TMF621 Trouble Ticket Management API, its resource model (ticket type, severity, status lifecycle, related entities and notes) is a practical normalization target across carriers [5]. TMF621 v5.0.0 is listed as the stable release as of October 2026 [5]. Expect custom fields for line tests and cause codes regardless, because those vary by vendor and access technology.
Illustrative telecom ticket record
The record below shows the structure that supports triage, guided troubleshooting and dispatch-avoidance modeling in one row set.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"ticket_id": "tkt_7f3a91",
"account_token": "acct_c2e4",
"circuit_token": "ckt_91bd",
"service_type": "DOCSIS_3.1",
"opened_at": "2025-03-14T18:02:00Z",
"channel": "chat",
"symptom_code": "INTERMITTENT_DROP",
"customer_text": "internet cuts out every evening for a few minutes",
"tests": [
{"t": "+00:04", "type": "modem_stats", "ds_power_dbmv": -9.8, "us_power_dbmv": 51.5, "t3_timeouts_24h": 37}
],
"actions": ["remote_reboot", "check_splitters_script"],
"escalation_tier": 2,
"parent_incident": null,
"dispatch": {"truck_roll": true, "outcome": "fixed", "tech_type": "drop"},
"cause_code": "DROP_CONNECTOR_CORROSION",
"resolution_code": "REPLACED_DROP_CONNECTOR",
"fault_location": "drop",
"repeat_within_30d": false
}
Note the tokenized account and circuit fields. They keep multi-ticket histories joinable without exposing the underlying identifiers.
CPNI and privacy limits on licensing ticket data
Trouble tickets from a US carrier can contain CPNI, so a release usually needs de-identification or aggregation plus counsel's sign-off. Under 47 U.S.C. § 222(h)(1), CPNI covers information about the quantity, technical configuration, type, destination, location and amount of use of a telecommunications service that a carrier obtains through the customer relationship [3]. Line configuration, test results and service location in a ticket fit that description closely.
The FCC's implementing rules in 47 CFR Part 64, Subpart U adopt the statutory definition [4]. The statute separately defines aggregate customer information as group data with individual identities and characteristics removed, which carriers may use and disclose more freely under § 222(c)(3) [3]. Whether a de-identified, record-level ticket file qualifies as aggregate information, and whether broadband-only tickets fall under § 222 at all given their regulatory classification, are questions for counsel rather than for the data team.
State privacy law applies on top. The CCPA's definition of deidentified information requires the holder to take reasonable measures against re-linking, publicly commit not to re-identify, and contractually bind recipients to the same [7]. NIST SP 800-188 is a sound reference for the method itself and warns that traditional de-identification has inherent limits [6].
Practical de-identification steps for tickets:
- Replace account numbers, phone numbers, emails and names with consistent tokens, scoped to the dataset.
- Tokenize circuit IDs, IP addresses, device MACs and CPE serials consistently so a repeat ticket or a 40-ticket outage cluster stays linked.
- Generalize service addresses to a wire center, node or ZIP3 level, and drop GPS coordinates from dispatch records.
- Scrub free-text notes, where agents paste callback numbers, gate codes and Wi-Fi passwords.
- Record the method and check a sample, because no method removes all re-identification risk.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
How reliable telecom cause codes are as labels
Cause and resolution codes are useful labels only after you measure how often they are wrong. Agents often pick a code at close under handle-time pressure, so "CPE replaced" can hide a drop fault that recurs, and "no trouble found" can mean the test ran after the fault cleared.
Ask suppliers for three checks before treating codes as ground truth. First, the repeat-ticket rate per cause code, since a code followed by a repeat within 30 days is suspect. Second, any post-resolution audit or QA sample comparing the code to technician notes. Third, code taxonomy changes over time, because a code-list merge or split silently changes what a label means. The guide to verifying outcome labels in operational records covers agreement checks in more depth.
Network-side root cause usually lives elsewhere. If your agent must reason from alarms to customer impact, pair tickets with NOC alarm and incident logs through the parent-incident link.
Request checklist for a telecom trouble ticket dataset
A precise request names the access technology, the fields, the label checks and the privacy treatment up front. Use this checklist when describing the data to a supplier or broker.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Item | What to specify | Why it matters |
|---|---|---|
| Access technology | FTTH, DOCSIS, DSL, fixed wireless, mobile, business data | Test fields and fault patterns differ by technology |
| Time span and volume | Months of history, ticket count, share with tests attached | Seasonality, storms and firmware rollouts shift distributions |
| Test evidence | Which automated tests, units and timing relative to open | Agents must learn from evidence, not just symptoms |
| Action trace | Ordered steps with timestamps and outcomes | Needed for guided-troubleshooting SFT and eval |
| Dispatch fields | Truck roll flag, outcome, repeat dispatch | Labels for dispatch-avoidance prediction |
| Label quality | Repeat rates by code, QA samples, taxonomy history | Determines whether codes can be training targets |
| Linkage | Parent incident IDs, consistent tokens across tables | Keeps outages and repeat histories intact |
| Privacy treatment | CPNI review, tokenization method, text scrubbing | Required before any release |
| Documentation | Data card covering source system, coding, known gaps | Supports audit and reuse [8] |
| Allowed uses | Training, evaluation, fine-tuning, term | Must match how the model will ship |
For enterprise IT incidents rather than carrier faults, the ITSM ticket datasets page is the better starting point. For general customer service conversations, see customer support ticket datasets.
Where SourceX fits for telecom ticket data
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Nothing is held in stock, a request does not guarantee a match, and every release is approved by the supplying company. You describe the data you need on the SourceX buyer intake; SourceX looks for US businesses that hold it.
Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded, and a sample is checked, though no method is perfect. The process runs Find, Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees.
For more context, see the telecom services buyer overview, what AI companies build with telecom data, data licensing rules for telecom companies, the industry-specific operational data hub, and the broader AI data buyer guides.
Request telecom trouble ticket data
If you are building ticket triage, guided troubleshooting or dispatch-avoidance models, describe the access technology, fields and allowed uses you need. SourceX looks for US businesses that hold matching data, reviews rights, and handles the license if a supplier agrees. Start at sourcex.si/buyers.
Sources
- arXiv (Sierra Research), "tau2-bench: A Dual-Control Benchmark for Agentic AI" (2025). https://arxiv.org/pdf/2506.07982
- arXiv (Sierra Research), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- Legal Information Institute, Cornell Law School, "47 U.S. Code 222 - Privacy of customer information". https://www.law.cornell.edu/uscode/text/47/222
- Legal Information Institute, Cornell Law School, "47 CFR 64.2003 - Definitions". https://www.law.cornell.edu/cfr/text/47/64.2003
- TM Forum, "Trouble Ticket Management API TMF621 v5.0" (2024). https://tmforum.org/oda/open-apis/directory/trouble-ticket-management-api-TMF621/v5.0
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- California Legislative Information, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.