Agent, workflow and domain-reasoning data
Runbook execution records for incident-response and IT operations agents
Quick answer
Runbook execution data for SRE agents is the linked record of an alert, the runbook a responder or automation invoked, each command or automation step actually run with its output, and the incident's resolution. Agents need the linkage, not runbooks alone: written procedures show intent, while execution logs show what worked, what was skipped and what failed. Source it from incident timelines, chat-ops logs and automation run histories, then scrub hostnames, IPs, credentials and vulnerability details before delivery.
By SourceX Editorial · Updated
Why runbook text alone does not train an operations agent
Runbook text alone teaches an agent what a team meant to do, not what it did under pressure. Current AI SRE products already treat runbooks as inputs: PagerDuty's SRE agent pulls runbooks from Confluence or GitHub alongside observability data as context [2], and Azure SRE Agent executes the diagnostic steps of a markdown runbook [3]. Practitioners also note that many runbooks are written for a human who fills gaps with tribal knowledge, so an agent following them literally stalls or takes the wrong branch [4].
Execution records close that gap. They show which steps were skipped because a dashboard already answered the question, which commands were re-run with different flags, and which escalations happened when a step failed. That is the signal you need for supervised fine-tuning on next-action prediction and for evaluation sets that score an agent against the path a resolved incident actually took. For the broader framing of agent data types, start at the agent training data hub.
The linkage that makes a record usable
A usable record joins five objects on a shared incident identifier: the triggering alert, the runbook version invoked, the ordered execution steps, the observed results and the resolution. Harness AI SRE illustrates the shape: it logs every runbook step with inputs, outputs and a running, success or failed status tied to the incident timeline [1]. If your supplier's tooling produces something similar, most of the join work is already done.
Where it does not, the same linkage has to be rebuilt from separate systems. Alerts come from the monitoring or paging tool, steps from automation engines and chat-ops bots, and the narrative from the ITSM record. In ServiceNow, for example, work notes and comments sit in sys_journal_field rather than on the incident row itself [7], so an export of the incident table alone loses the timeline. The cross-system workflow records guide covers joining ITSM, chat and email into one task timeline.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"incident_id": "INC-REDACTED-0193",
"alert": {
"source": "monitoring",
"signal": "p99_latency_ms > 1200 for 10m",
"service": "svc-checkout",
"fired_at": "2025-03-14T02:11:05Z"
},
"runbook": {
"id": "rb-checkout-latency",
"version": "git:3f9c1a2",
"format": "markdown"
},
"steps": [
{"seq": 1, "actor": "human:oncall_A", "channel": "chatops",
"action": "kubectl get pods -n <ns> -l app=<svc>",
"status": "success", "output_summary": "2 of 6 pods CrashLoopBackOff"},
{"seq": 2, "actor": "automation", "channel": "runbook_engine",
"action": "restart_deployment", "status": "failed",
"output_summary": "readiness probe timeout"},
{"seq": 3, "actor": "human:oncall_A", "channel": "chatops",
"action": "rollback to previous release tag", "deviation_from_runbook": true,
"status": "success"}
],
"resolution": {
"resolved_at": "2025-03-14T02:48:40Z",
"outcome": "mitigated_by_rollback",
"root_cause_category": "bad_deploy",
"postmortem_ref": "PM-REDACTED-041"
}
}
Two fields matter more than they look. deviation_from_runbook marks where the responder left the documented path, which is often the most valuable training signal. runbook.version lets you check that the procedure an agent learns from is the one that was live at the time, not a later edit.
Where execution records live in a supplying company
Execution records are spread across four or five systems in most operations teams, and each has its own retention and redaction profile. Knowing where they live helps you write a request a supplier can actually answer.
| Source system type | What it contributes | Typical gaps | Redaction focus |
|---|---|---|---|
| Paging and alerting (on-call tools) | Alert payload, acknowledgement, escalation chain, timestamps | Little detail on actions taken | Responder names, phone numbers |
| Runbook automation engines | Step inputs, outputs, status per step [1] | Only covers automated steps | Hostnames, IPs, API tokens in parameters |
| Chat-ops channels | Commands typed, bot responses, human reasoning in thread | Unstructured, interleaved incidents | Pasted secrets, customer identifiers |
| ITSM incidents and work notes | Priority, assignment, work notes [7], closure codes | Notes written after the fact | Employee and customer names |
| Postmortems | Root cause, contributing factors, follow-up actions | One per major incident only | Vulnerability details, vendor names |
Postmortems and ITSM tickets have their own owner pages: see incident postmortems and ITSM ticket datasets. This page covers the runbook-to-execution linkage between them. Infrastructure code and CI logs belong with code data, not here. Endpoint alerts handled by managed service providers are covered in the MSP RMM alert-to-remediation records guide.
Labeling outcomes and deviations for SFT and evaluation
Outcome labels should come from the business record, not from an annotator's guess. Use closure codes, time to mitigation and whether the incident reopened within a window you define. The task success labels guide explains how to derive those definitions from operational records.
For supervised fine-tuning, frame each step as a state-action pair: the state is the alert, the runbook text and prior step outputs; the action is the next command or decision. For evaluation, hold out whole incidents and score an agent on whether it reaches the same mitigation, how many steps it takes, and whether it attempts actions the responder ruled out. If your agent emits OpenTelemetry traces, the execute_tool span conventions, where tool call arguments and results are opt-in attributes [6], give you a format for comparing agent runs against historical steps.
Watch for three failure modes in the labels:
- Survivorship bias: only incidents with clean automation logs get exported, so the hard, manual ones are missing.
- Retroactive notes: work notes typed after resolution read as reasoning but were written knowing the answer.
- Runbook drift: the runbook was edited after the incident, so the text and the steps disagree.
The exception handling records guide covers the long-tail cases that clean runbooks rarely capture.
Security redaction before anything leaves the supplier
Operations records carry more security-sensitive material than most business data, so redaction scope must be agreed before sampling. Hostnames, internal IP ranges, cloud account IDs, API keys, tokens pasted into chat and unpatched vulnerability details can all appear in step outputs. Models can reproduce verbatim training sequences, including identifiers such as UUIDs [8], so a secret left in a command log is a real leak risk in a trained model, not only in the dataset.
A practical approach is consistent pseudonymization of infrastructure identifiers (so svc-checkout stays the same token across an incident), secret scanning on every free-text field, and generalization of CVE-specific remediation into categories where the supplier requires it. Security-incident records also follow different handling norms; NIST SP 800-61 Rev. 3 frames incident response inside broader cybersecurity risk management [5], and DFIR case data is covered separately in the incident response reports guide.
Request checklist for runbook execution data
A precise request names the linkage, fields, time window and redaction rules up front. Use this as a starting template.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Scope: incident types (latency, capacity, deploy failures, certificate expiry), severity levels and services in scope
- Linkage: each record joins alert, runbook ID and version, ordered steps and resolution on one incident key
- Step fields: actor type (human, bot, automation), channel, action text or command, status, output summary, timestamp
- Deviation flag: whether each step followed, skipped or departed from the runbook
- Outcome: closure code, time to mitigation, reopen flag, postmortem reference where one exists
- Redaction: hostnames, IPs, account IDs, secrets, personal names and vulnerability specifics, with the method documented
- Coverage evidence: share of incidents with full step logs versus partial, to expose survivorship bias
- Format: JSONL per incident, plus runbook texts as markdown keyed by version
- Allowed use: training, evaluation or both, stated in the license
If you are weighing licensed records against commissioned demonstrations or synthetic incidents, the licensed, commissioned or synthetic trajectories comparison sets out the tradeoffs.
How SourceX handles requests for runbook execution data
SourceX sources operational datasets from US companies on request; it does not hold runbook data in stock, and a request does not guarantee a match. Engineering records are among the kinds of data it sources. Buyers describe the data they need, SourceX looks for US businesses that hold it, and every release is approved by the supplying company. You can describe the records you need on the buyers page.
Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails and phone numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. For the capability view, see training data for IT operations agents and what IT incident recovery records show.
Request runbook execution data for your SRE agent
SourceX serves AI teams wherever they are based and manages the commercial process from finding a supplier through licensing and ongoing purchases. Nothing is contracted until a supplier agrees, and terms are agreed per deal. Start a runbook execution data request.
Sources
- Harness, "How to build runbooks that work, and automate them with Harness AI SRE". https://www.harness.io/blog/how-to-build-runbooks-that-work----and-automate-them-with-harness-ai-sre
- PagerDuty Engineering, "Context over cleverness: building PagerDuty's SRE agent". https://www.pagerduty.com/eng/context-over-cleverness-building-pagerdutys-sre-agent/
- Microsoft Tech Community, "Azure SRE Agent: Automate Runbook Execution with GenAI". https://techcommunity.microsoft.com/blog/appsonazureblog/-/4479811
- tianpan.co, "The agent runbook your incident commander could not execute" (2026). https://tianpan.co/blog/2026/06/02/the-agent-runbook-your-incident-commander-could-not-execute
- NIST, "NIST Revises SP 800-61: Incident Response Recommendations and Considerations for Cybersecurity Risk Management". https://content.govdelivery.com/accounts/USNIST/bulletins/3d9dba1
- OpenTelemetry, "Semantic conventions for generative client AI spans". https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-spans
- ServiceNow, "Journal fields (ServiceNow Documentation)". https://docs.servicenow.com/bundle/vancouver-platform-administration/page/administer/field-administration/concept/c_JournalFields.html
- USENIX Security 2021 (Carlini et al.), "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.