Industry-specific operational data
Utility Outage Tickets and Restoration Logs for AI
Quick answer
Utility outage data for machine learning means event-level records exported from an outage management system (OMS): one row per interruption with start and restore timestamps, the operating device, feeder and substation, customers interrupted, customer minutes interrupted, cause and sub-cause codes, and crew steps. Public outage maps and reliability statistics are summaries, not tickets. Restoration-time and cause models need the ticket history itself, licensed from the utility or contractor that holds it, with customer data and sensitive grid detail handled before delivery.
By SourceX Editorial · Updated
What an OMS outage ticket contains
An OMS outage event is a structured record built from customer calls, AMI last-gasp messages and SCADA trips, then edited by dispatchers and crews until the event closes. The fields that matter for modeling sit in three layers: the event header, the crew activity log, and the free-text remarks. Systems such as OMS modules in ADMS suites, or homegrown trouble-ticket tools at co-ops and municipals, store these in different tables, so an export is usually a join across event, device, crew and call tables.
Expect these fields in a usable extract:
- Event header: event ID, parent/child (nested) event links, first-call time, outage start, restore time (often per step for partial restorations), event status and close time.
- Location in the network: operating device type (fuse, recloser, breaker, switch, transformer), device ID, feeder, substation, phase, and a coarse geography such as service territory zone.
- Impact: customers interrupted (CI), customer minutes interrupted (CMI), and whether the event was momentary or sustained.
- Cause: cause code and sub-cause (vegetation, animal, equipment failure, lightning, vehicle, unknown), weather flag, and a major event day indicator.
- Crew actions: crew assignment, dispatch, arrival, switching steps, ETR (estimated time of restoration) revisions and the final restoration step.
- Remarks: short crew notes such as "tree on line," "squirrel on xfmr," or "blown fuse, refused," which are the language layer for cause classification and dispatcher agents.
How IEEE 1366 and major event days shape the labels
IEEE 1366 decides which events count toward SAIDI, SAIFI and CAIDI, so it shapes how utilities code, review and sometimes exclude events. The standard separates major events from underlying reliability trends by flagging any day whose daily system SAIDI exceeds a statistical threshold, usually labeled TMED, derived from several prior years of daily SAIDI. Many state commissions and federal reporting programs ask for reliability both with and without these major event days, so the flag travels with the data.
For a buyer, this has three consequences. Storm days are often coded more loosely because crews prioritize restoration over documentation, so cause codes on major event days carry more "unknown" and back-filled values. Exclusion flags can be mistaken for data-quality filters and silently drop the very storm events a storm-restoration model needs. And restoration times inside nested events may be reported per customer group, so a single "restore time" column understates long tails.
Ask the supplier which edition of IEEE 1366 they follow, how they compute TMED, and whether the export includes MED-flagged days. Also ask whether multi-day interruptions are accrued to the day they began, since that convention changes which events fall inside a storm window.
Public outage data versus licensed ticket history
Public sources are good for weather joins and benchmarking, but they rarely give event-level tickets with cause codes and crew steps. Argonne's work, for example, built a platform from customer outage announcements on public utility websites and emphasized that ML results depend on data quality and quantity [1]. Published research that does use event-level OMS records typically gets them through utility partnerships [2][3].
| Need | Public outage feeds and reliability reports | Licensed OMS ticket history |
|---|---|---|
| Restoration-time (ETR) prediction | Customer counts over time, no crew steps | Start, ETR revisions, restore per step |
| Cause classification | Rarely includes causes | Cause and sub-cause codes plus crew remarks |
| Storm outage forecasting | County-level counts | Device and feeder-level events with MED flags |
| Dispatcher agent evaluation | Not usable | Event timelines, switching steps, crew assignments |
Storm planning studies combine outage history with weather to forecast damage and stage crews [4]; vegetation-outage models do the same with vegetation and asset attributes [2]. If your model also needs tree-work history, see utility vegetation management records.
Joining weather, assets and time correctly
Outage tickets become predictive only when joined to weather and asset attributes at the time of the event. Useful joins include hourly wind gust and precipitation, lightning strikes, conductor type, pole and device age, and vegetation cycle dates. Each join must be as-of the event start, not the current asset register, or a model learns from equipment replaced after the outage; see point-in-time correct training data.
Split train and test by time and by storm, never randomly by row. Child events from one storm leak across a random split and inflate restoration-time accuracy. Define ground truth for ETR as the final restore time per customer group, and keep each ETR revision as a separate timestamped field so you can measure how early predictions were.
Privacy and grid-security handling before delivery
Outage records mix customer data with network detail, and both need treatment before training. Customer-level fields include account numbers, service addresses, phone numbers from call records, and flags such as medical-baseline or life-support customers, which should be removed or replaced. Location should be generalized to feeder or grid cell where your model allows it.
Grid topology, critical-facility locations and switching plans can be security-sensitive. Federal rules on Critical Energy/Electric Infrastructure Information (CEII) restrict how designated grid information is shared, and utilities often extend similar caution to internal network detail. Utility ticket exports are not automatically CEII, but suppliers may treat one-line diagrams, substation identities and feeder maps conservatively, so plan for device IDs to be tokenized and substation names masked. Confirm with counsel whether any field you request has been designated.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Request template for outage and restoration records
A precise request describes the records and the model, not the utility. Use a scoping template like the one below so a supplier can confirm whether its OMS holds the data.
Illustrative example: invented to show structure; it does not describe an available dataset.
request: utility_outage_events
use_case: restoration-time prediction and cause classification
systems: OMS event tables, crew activity log, call records (counts only)
grain: one row per outage event, plus one row per restoration step
history: multiple years including at least several major storms
fields_required:
- event_id, parent_event_id
- outage_start, first_call_time, restore_time, etr_revisions[]
- device_type, device_id (tokenized), feeder_id (tokenized), substation (masked)
- customers_interrupted, customer_minutes_interrupted
- cause_code, sub_cause_code, weather_flag, med_flag
- crew_dispatch_time, crew_arrival_time, switching_steps
- crew_remarks (free text, scrubbed)
excluded: account numbers, service addresses, phone numbers, medical-baseline flags
documentation: cause code dictionary, IEEE 1366 edition and TMED method, code changes by year
quality_checks: share of unknown causes, restore_time < outage_start, duplicate events
Ask for the cause-code dictionary and its revision history; a recode in a single year can look like a reliability trend. Measure label quality against a defined process, such as the ISO/IEC 5259-4 data quality framework for supervised labelling [5]. Related operational data with similar structure includes telecom network trouble tickets and PLC and SCADA alarm logs.
How SourceX handles outage data requests
SourceX sources operational datasets from US companies on request, including engineering records and documents; it does not hold outage data in stock, and a request does not guarantee a match. You describe the records you need, and SourceX looks for US businesses that hold them, with every release approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Start a scoped request on the SourceX buyer page, or compare adjacent categories such as dispatch logs, maintenance work orders and geospatial and location data. For the wider landscape, see the industry-specific operational data guide and the AI data hub.
Request utility outage records for your model
If you are building restoration-time, storm or cause models, describe the OMS fields, history and use you need. SourceX runs Find, Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Describe your outage data request.
Sources
- Argonne National Laboratory, "Outage prediction and grid vulnerability identification using machine learning on utility outage data". https://www.anl.gov/esia/outage-prediction-and-grid-vulnerability-identification-using-machine-learning-on-utility-outage
- arXiv, "A Data-Driven Approach for Predicting Vegetation-Related Outages in Power Distribution Systems" (2018). https://ar5iv.arxiv.org/html/1807.06180
- Southern Methodist University, "SMU Data Science Review, Volume 6, Issue 1, Article 5". https://scholar.smu.edu/datasciencereview/vol6/iss1/5
- SciTePress, "Data Analytics for Power Utility Storm Planning" (2014). https://www.scitepress.org/Papers/2014/51282/51282.pdf
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Data quality for analytics and ML - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.