Tables, time series and transactional data
Telemetry Aligned with Incident Labels for Root-Cause Analysis Models
Quick answer
Root-cause analysis (RCA) models need logs, metrics and traces from the same window, joined to an incident record that states when the failure began and ended, which services were impacted, and which component was confirmed as the cause. Public sets are mostly lab fault injections or anomaly-flagged logs, so production-grade RCA data usually has to be licensed from operators. Specify the label schema, clock-alignment rules and evaluation split before you request anything.
By SourceX Editorial · Updated
Why public RCA benchmarks run out quickly
Public telemetry with confirmed root causes is scarce, and what exists mostly comes from injected faults in demo applications. Loghub, the most cited public log collection, covers many systems, but only a minority of its datasets are labeled and those labels mostly mark anomalous lines or sessions rather than confirmed root causes [1]. Recent LLM log-analysis benchmarks still recycle a handful of these public sets, with no multi-source root-cause ground truth [2].
Academic multi-source RCA benchmarks exist, typically built by injecting faults into open demo microservice applications such as Online Boutique, Sock Shop and Train Ticket and recording metrics, logs and traces around each fault [5]. They are good harnesses for comparing methods, but the faults are usually injected one at a time (CPU hogs, memory leaks, network delay, code-level bugs) into small topologies, and the label is known by construction rather than confirmed by responders.
Production incidents look different. They involve cascading failures across dozens or hundreds of services, partial deploys, config pushes, third-party dependency outages, noisy alerting, and missing telemetry exactly where the fault occurred. A model that tops a fault-injection leaderboard can still fail on change-induced incidents, which is why teams building SRE agents look for licensed operational data. The broader landscape of tabular and time-series sources is mapped in our buyer's guide to tabular, time-series and transactional data.
What a usable incident-labeled telemetry record contains
A usable record is an incident window, not a log dump: telemetry for a bounded time range plus a structured label that a responder or reviewer confirmed. The label usually comes from the incident management system (PagerDuty, Opsgenie, ServiceNow, Jira Service Management) and the postmortem, and it should be normalized into fields rather than left as prose.
Minimum label fields worth requiring:
- Incident timeline: detection time, declared start, customer-impact start and end, mitigation time and resolution time, each in UTC with its source system.
- Impact scope: impacted services, endpoints, regions and SLOs, keyed to the same service names used in the telemetry.
- Root-cause component: the service, host, pod, database, dependency or config object confirmed as cause, with a granularity flag (service, instance, metric, log template or code change).
- Trigger and change link: deploy ID, feature-flag change, config commit or infrastructure event, if any.
- Remediation: rollback, restart, scale-out, failover or hotfix, and who confirmed it.
- Label confidence: confirmed by postmortem, inferred by responder, or disputed.
Telemetry should arrive in columnar or line-delimited formats with stable schemas: metrics as long-format time series (timestamp, metric name, labels, value), logs with raw text plus parsed template IDs, and traces as span tables with trace ID, span ID, parent span ID, service, operation, start, duration and status. Parquet is a practical default for the span and metric tables because its footer metadata lets loaders read only the needed column chunks [3]. Many of the same log-handling questions appear in our guide to production log data for anomaly detection.
Illustrative example: invented to show structure; it does not describe an available dataset.
incident_id: INC-0417
label_source: postmortem_reviewed # postmortem_reviewed | responder_note | inferred
timeline_utc:
detected: 2025-03-11T14:07:42Z
impact_start: 2025-03-11T13:58:10Z
mitigated: 2025-03-11T14:41:05Z
resolved: 2025-03-11T16:02:00Z
impacted_services: [checkout-api, payment-gateway]
root_cause:
component: payment-db-replica-2
granularity: instance # service | instance | metric | log_template | code_change
category: resource_exhaustion/connection_pool
trigger: config_change
change_ref: cfg-commit-8c1e
remediation: config_rollback
telemetry_window_utc: [2025-03-11T13:00:00Z, 2025-03-11T17:00:00Z]
telemetry:
metrics: metrics/INC-0417.parquet # ts, metric, labels(map), value
logs: logs/INC-0417.parquet # ts, service, host, level, template_id, message_redacted
traces: traces/INC-0417.parquet # trace_id, span_id, parent_span_id, service, op, start, dur_us, status
clock_notes: "app hosts NTP-synced; load balancer logs local time UTC-5, converted; max observed skew 180 ms"
annotator_agreement: {category_kappa: 0.71, reviewers: 2}
Aligning clocks across logs, metrics and traces
Alignment errors silently destroy RCA labels, so require the supplier to document every timestamp's source, timezone and known skew. A two-second offset between a metrics scraper and an application log can reverse the apparent order of cause and symptom, and order is exactly what RCA models learn.
Ask for these alignment facts per source:
- Timestamp field name, precision (seconds, milliseconds, microseconds) and whether it records event time or ingestion time.
- Timezone handling, including any local-time device or load-balancer logs and daylight-saving transitions inside the window.
- Clock synchronization method (NTP, PTP, cloud time service) and any measured skew.
- Metric scrape or aggregation interval, since a 60-second rollup blurs onset by up to a full interval.
- Sampling: head- or tail-based trace sampling rates and log sampling or rate limits during the incident.
Trace propagation also matters. Spans propagated with the W3C Trace Context traceparent header carry a consistent trace identifier across services, which lets you join spans from different vendors and runtimes into one causal graph. If some services do not propagate context, the call graph will have gaps, and the supplier should mark which edges are inferred from logs rather than observed in traces. The same as-of discipline described in our guide to point-in-time correct training data applies: features for a prediction at time t must use only telemetry ingested before t.
Building a root-cause taxonomy and checking label agreement
Root-cause labels are only trainable if they map to a fixed taxonomy and two reviewers would assign the same category. Free-text postmortem causes such as "bad deploy" or "DB issues" cannot be scored.
A workable two-level taxonomy separates the faulty component type (application code, configuration, dependency, data store, network, compute, platform) from the failure mechanism (resource exhaustion, latency injection, error-rate spike, crash loop, data corruption, capacity, expired certificate). Ask the supplier, or your own annotators, to double-label a sample and report Cohen's or Fleiss' kappa per category. Low agreement on a category is a signal to merge it, not a reason to drop the incidents.
Keep postmortem narratives linked but separate. The text belongs to a different data product, covered on our page for licensing incident postmortems, and pairing it with telemetry by incident ID enables multimodal RCA where a model reads both the signals and the responder's reasoning. For the reverse direction, telemetry paired with operator text, see paired time-series and text data.
Evaluating RCA and SRE agents on licensed incidents
Evaluate on held-out incidents with metrics that reward ranking the true cause highly and finding it fast. Many RCA papers report top-k accuracy (whether the confirmed root cause appears among the model's top k candidates) and an average across k, which you can reproduce on licensed incidents if root-cause labels use one consistent granularity.
Add metrics that matter in production:
| Metric | What it measures | Data it requires |
|---|---|---|
| Top-1 / top-3 / top-5 accuracy | Ranked candidate list contains the confirmed cause | Root-cause component at a consistent granularity |
| Time-to-localize | Minutes from impact start until the model's top candidate is correct, using only data available at each step | Event-time stamps plus ingestion lag |
| Mechanism accuracy | Correct failure category, not just the component | Taxonomy labels with agreement scores |
| Change attribution | Correct deploy or config change identified | Change log joined by service and time |
| False-alarm cost | Candidates that would trigger unnecessary rollbacks | Remediation field and negative windows |
Split by time, not randomly: train on earlier incidents and test on later ones so topology changes and new services appear in test. Also hold out entire services or incident classes to test transfer. Include quiet windows with no incident so agents learn to say "no root cause" and you can measure false positives. For agents that also open tickets or query runbooks, link incidents to records described on our ITSM ticket datasets page and our overview of training data for IT operations agents.
Diligence and redaction for operational telemetry
Telemetry carries more sensitive content than teams expect, so treat redaction and rights review as part of the specification. Log messages and span attributes routinely contain customer emails, IP addresses, account IDs, session tokens, internal hostnames and query parameters.
Practical requirements for any supplier:
- Redact or tokenize personal and secret values consistently, so the same customer ID maps to the same token across logs and traces and joins still work.
- Keep service and host names stable, or pseudonymize them with a published mapping, because RCA depends on topology.
- Provide a service dependency map or let you derive it from traces, and state the topology's effective date.
- Document the window selection rule (for example, a fixed lead time before impact start and a lag after resolution) so the dataset does not leak the label through window boundaries.
- Supply dataset-level metadata in a machine-readable form; Croissant JSON-LD describes files and record structure in a way common loaders understand [4].
SourceX sources operational datasets from US companies on request, including engineering records of this kind, and manages the commercial process through licensing and ongoing purchases. Nothing is held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery with the method recorded and a sample checked, and no method is perfect. You can describe the telemetry you need to SourceX, and the supplying company approves every release.
Request template for incident-labeled telemetry
A precise request describes the data and labels, not the companies that might hold them. Use this as a starting point.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example specification |
|---|---|
| Use | Train and evaluate an RCA ranking model and an SRE agent |
| System type | Microservice platform on Kubernetes, 50+ services, OpenTelemetry tracing |
| Signals | Metrics (15–60 s resolution), parsed logs, sampled traces, deploy and config change log |
| Label source | Incident management records plus reviewed postmortems |
| Label fields | Timeline (UTC), impacted services, root-cause component and granularity, category, trigger, remediation, confidence |
| Window rule | 60 min before impact start to 60 min after resolution, plus matched quiet windows |
| Alignment | Documented timezones, sync method, observed skew, sampling rates |
| Redaction | Consistent tokenization of personal and secret values; stable service names |
| Formats | Parquet per signal per incident; YAML or JSON label file; Croissant metadata |
| Evaluation split | Time-based holdout of the most recent incidents |
Related structures appear in our guides to cross-system workflow records and network alarm and incident logs for root-cause AI, which covers telecom NOC data rather than application telemetry.
Source incident-labeled telemetry for RCA with SourceX
SourceX looks for US businesses that hold the telemetry and incident records you describe, assesses the data and its licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Start by describing your RCA data requirements on the buyers page.
Sources
- arXiv (He et al.), "Loghub: A Large Collection of System Log Datasets for AI-driven Log Analytics" (2020). https://arxiv.org/pdf/2008.06448
- arXiv, "LLM-Enhanced Log Anomaly Detection: A Comprehensive Benchmark of Large Language Models for Automated System Diagnostics" (2026). https://arxiv.org/pdf/2604.12218
- The Apache Software Foundation (Apache Parquet), "File Format". https://parquet.apache.org/docs/file-format/
- arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- arXiv (Pham et al.), "RCAEval: A Benchmark for Root Cause Analysis of Microservice Systems with Telemetry Data" (2025). https://arxiv.org/abs/2412.17015
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.