Industry-specific operational data
NOC alarm and incident logs for network root-cause AI
Quick answer
Root-cause models for network operations need more than log lines: they need alarm streams joined to topology snapshots, correlated incident groups and the ticket or postmortem that names the confirmed cause. Public log benchmarks such as Loghub's HDFS and BGL come from software systems and supercomputers, not carrier networks, and rarely carry root-cause labels. When licensing NOC data, require normalized alarm fields, raise and clear times, maintenance-window flags, a ticket join with a confirmed root cause, and tokenized element names.
By SourceX Editorial · Updated
Why public log benchmarks do not cover network alarm correlation
Public log datasets are useful for parsing and generic anomaly detection, but they were not collected from network element managers and contain no alarm correlation ground truth. Loghub, the main open collection, provides 19 real-world log datasets from distributed systems, supercomputers and operating systems, and only a minority are labeled [1]. None of them is a carrier alarm stream with topology.
The gap persists in newer work. A 2026 benchmark of LLM-based log anomaly detection still evaluates on a handful of classic datasets, including HDFS, BGL, Thunderbird and Spirit [2], and the LogEval suite reuses the same public logs across parsing and anomaly tasks [3]. Results on those corpora say little about how a model handles a fiber cut that raises hundreds of loss-of-signal alarms across a ring.
Network alarm data differs in three structural ways:
- Stateful, not line-oriented. An alarm is raised, possibly updated in severity, and cleared. Modern alarm models treat alarms as states on resources rather than discrete notifications, so a row-per-line log representation loses the lifecycle.
- Topology-dependent. Whether two alarms share a cause depends on adjacency, shared links, shared power or shared software releases. Without a topology snapshot at the time of the incident, correlation labels cannot be learned or checked.
- Label lives elsewhere. The confirmed cause is usually written in a trouble ticket, a change record or a postmortem, not in the alarm stream itself.
Teams working on software telemetry should compare this page with telemetry aligned with incident labels for root-cause analysis and production log data for log anomaly detection; NOC data is a separate type with its own normalization problems.
What a labeled NOC alarm and incident record contains
A usable record joins four layers: the raw alarm event, its normalized fields, the correlation group, and the linked incident with a confirmed root cause. Each layer comes from a different system, typically an element management system (EMS) or network management system (NMS) for alarms, an inventory or topology database, a fault-management correlation engine, and an ITSM tool for tickets.
Published standards make normalization feasible across vendors; confirm the exact attribute definitions against the current texts. ITU-T X.733 [4] defines the perceived severity scale (critical, major, minor, warning, indeterminate, cleared) and probable cause attributes, and IETF RFC 5674 [5] carries those values into syslog with an enumerated probable cause such as transmissionError. TM Forum's alarm data model groups probable causes under event types such as communications, equipment, environmental, quality of service and processing error alarms, with values consistent with X.733 or 3GPP TS 32.111-2 Annex B. RFC 8632 [6] adds a YANG alarm model with an alarm inventory of every alarm type a device can raise and explicit shelving of suppressed alarms.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value | Why it matters for RCA models |
|---|---|---|
| alarm_id | A-20260314-0042817 | Stable key for joins and deduplication |
| source_element | NE_7f3a91 (tokenized) | Node identity without exposing real hostnames |
| element_type | OTN transponder | Lets models generalize across vendors |
| event_type | communicationsAlarm | X.733 category for normalization |
| probable_cause | lossOfSignal | Enumerated cause, not free text |
| specific_problem | vendor code LOS-OCH | Vendor detail preserved for later mapping |
| perceived_severity | critical | Severity at raise; track changes separately |
| raised_at / cleared_at | 2026-03-14T02:11:07Z / 02:49:55Z | Duration and flapping detection |
| correlation_group_id | CG-55120 | Ground truth or engine-assigned grouping (must be labeled which) |
| is_root_alarm | true | The alarm the operator judged causal |
| maintenance_window | false | Separates planned work from faults |
| shelved | false | Suppressed alarms distort base rates |
| ticket_id | INC-88213 | Join to the incident or trouble ticket |
| confirmed_root_cause | fiber cut, span 12, third-party excavation | The training label; source and reviewer recorded |
| topology_snapshot_id | TOPO-2026-03-14T00 | Graph state at incident time |
The correlation_group_id field needs a provenance flag. A group created by the operator's existing correlation engine teaches a model to imitate that engine; a group confirmed by an engineer after resolution teaches the real relationship. Ask for both and keep them distinct.
Where confirmed root causes actually come from
Confirmed root causes almost always come from the incident record, not the alarm feed, so the dataset is only as good as the alarm-to-ticket join. In most NOCs a ticket is opened from one alarm or a correlation group, worked by tier 1 and tier 2 staff, and closed with a resolution code and notes. Major incidents get a postmortem with a written timeline and causal chain.
Common failure modes buyers should check for:
- Resolution-code drift. Closure codes such as "hardware" or "no fault found" are chosen under time pressure. Sample tickets and compare the code with the narrative notes.
- Missing join keys. Tickets opened by phone or email may never reference an alarm ID. Measure the share of tickets with a valid alarm or group link before you agree on scope.
- Post-hoc correlation leakage. If the correlation group was edited after the ticket closed, features computed from it leak the answer into training.
- Change-induced faults. Many network incidents trace to a configuration change or software upgrade. Without change records or maintenance-window flags, a model learns that upgrades cause random alarms.
Related joins appear in telecom network trouble ticket data for troubleshooting agents, and SourceX also covers licensing incident postmortems and ITSM ticket datasets for adjacent use cases.
Building training and evaluation splits from alarm streams
Split by time and by incident, never by individual alarm, or the model will see sibling alarms from the same outage in both training and test sets. A single power failure at a hub site can produce thousands of alarms that share one cause; random row-level splits make correlation look trivially easy.
Practical split rules:
- Hold out whole calendar windows (for example, the final quarter of the export) as the test period.
- Keep every alarm in a correlation group on the same side of the split.
- Record topology snapshots per window, since network changes between training and test periods are part of the real difficulty.
- Report results separately for storms (large groups), single-alarm incidents and change-related incidents.
For NOC agent evaluation, build tasks from resolved incidents: give the agent the alarm window, topology and runbook, then score whether it names the root element and cause category that the ticket confirms. Time-aligned alarm, metric and note data is discussed further in paired time-series and text data.
Security and privacy limits on network data
Topology, element names, IP address plans and circuit IDs are sensitive infrastructure details, so expect suppliers to generalize or tokenize them and to restrict access. Hostnames often encode site codes and city names, and circuit IDs can reveal enterprise customers. Tokenize consistently so that the same element maps to the same token across alarms, topology and tickets, otherwise joins break.
Ticket notes carry personal data and customer details: technician names, callback numbers, customer account numbers and street addresses for site visits. SourceX removes or replaces personal details such as names, emails, phones and account numbers before delivery, records the method, and checks a sample, though no method is perfect. Rare events and small sites can still act as indirect identifiers, a risk covered in indirect identifiers in business text.
Requirements checklist for licensing NOC alarm data
A NOC data license should define the alarm schema, the label source, the join coverage and the security treatment before any data moves. Use this checklist when you write a request or review a supplier's sample.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Requirement | What to ask for | Red flag |
|---|---|---|
| Alarm normalization | Mapping from vendor codes to X.733 event type, probable cause and severity | Free-text alarm descriptions only |
| Lifecycle | Raise, severity change and clear events with timestamps in UTC | Only active alarms at export time |
| Suppression state | Shelved, acknowledged and maintenance flags | Suppressed alarms silently dropped |
| Correlation labels | Engine-assigned and engineer-confirmed groups, flagged separately | One unlabeled "group" column |
| Root-cause label | Ticket or postmortem cause, with closure code and notes | Closure codes with no narrative |
| Join coverage | Share of tickets linked to an alarm ID or group | No measurement offered |
| Topology | Time-stamped graph snapshots with tokenized nodes and link types | A single current-state topology |
| Change records | Change ID, window and affected elements | No link between changes and alarms |
| Tokenization | Consistent tokens across alarms, topology and tickets | Different hashing per table |
| Rights | Evidence the supplier can license the data and any vendor-format restrictions | Data exported from a vendor's managed service without its approval |
The rights row matters because alarm data may sit in a vendor-operated managed service or an outsourced NOC. Ask who owns the export and whether any contract limits reuse; see chain of title for AI training data.
How SourceX approaches NOC alarm and incident data
SourceX sources operational datasets from US companies on request and manages the licensing process; categories such as network alarm data are not held in stock, and a request does not guarantee a match. Buyers describe the data they need, and SourceX looks for US businesses that hold it, with every release approved by the supplying company.
Each dataset is reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Telecom buyers can also see buyers by industry: telecom services, training data for IT operations agents, and the wider industry-specific operational data guide. To describe a specific alarm and ticket scope, use the SourceX buyer request page.
Request network alarm data for root-cause AI
SourceX finds US companies that hold the operational data you describe, assesses data and licensing permissions, and manages the license and ongoing purchases; nothing is contracted until a supplier agrees. Describe the alarm, topology and ticket fields you need on the SourceX buyers page.
Frequently asked questions
Can SCADA or industrial alarm logs substitute for NOC data?
Only partly. Industrial alarm logs share the raise-clear lifecycle and alarm-flood problems, but the topology, protocols and cause taxonomy differ; see PLC, SCADA and DCS alarm and event logs for that data type.
Should we take raw syslog and SNMP traps or normalized alarms?
Ask for both when possible. Raw traps and syslog preserve vendor detail for parsing models, while normalized X.733-style records with lifecycle and correlation fields are what RCA and correlation models train on.
How much history is useful?
Enough to cover seasonal load, at least one major software upgrade cycle and a meaningful number of confirmed major incidents. Count confirmed root-cause incidents rather than raw alarms, since alarm volume is dominated by storms and flapping.
Sources
- arXiv (He et al.), "Loghub: A Large Collection of System Log Datasets for AI-driven Log Analytics" (2020). https://arxiv.org/pdf/2008.06448
- arXiv, "LLM-Enhanced Log Anomaly Detection: A Comprehensive Benchmark of Large Language Models for Automated System Diagnostics" (2026). https://arxiv.org/pdf/2604.12218
- arXiv, "LogEval: A Comprehensive Benchmark Suite for Large Language Models In Log Analysis" (2024). https://arxiv.org/pdf/2407.01896
- ITU-T, "X.733 : Information technology - Open Systems Interconnection - System Management: Alarm reporting function" (1992). https://www.itu.int/rec/T-REC-X.733-199202-I/en
- IETF, "RFC 5674: Alarms in Syslog" (2009). https://datatracker.ietf.org/doc/html/rfc5674
- IETF, "RFC 8632: A YANG Data Model for Alarm Management" (2019). https://datatracker.ietf.org/doc/html/rfc8632
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.