Skip to content

Industry-specific operational data

Labeled Equipment Failure Data for Predictive Maintenance Models

Quick answer

A usable predictive maintenance dataset with failure labels joins condition-monitoring signals (vibration, temperature, motor current, oil analysis) to failure events taken from corrective work orders, with failure mode, cause and remedy codes and dated detection and repair. Most public options are simulated or lab run-to-failure sets. Real fleet data is sparse in failures, censored and full of interventions, so buyers should specify asset class, signals, label provenance and labeled event counts per failure mode before evaluating any supplier.

By SourceX Editorial · Updated

Why public benchmarks rarely transfer to real fleets

Public predictive maintenance benchmarks are mostly synthetic or lab-generated, so models tuned on them learn degradation physics that real plants do not produce. The UCI AI4I 2020 dataset, one of the most cited results for this query, is a synthetic set of about 10,000 rows with a machine-failure flag and five failure modes: tool wear, heat dissipation, power, overstrain and random failures [1]. NASA's C-MAPSS turbofan data, the standard remaining useful life (RUL) benchmark, is likewise produced by a simulation model rather than by engines in service, with training units run to failure and test units truncated before it.

Lab bearing rigs share the same issue in a different form. They run a component under controlled load until it fails, with no operator who hears a noise and swaps the bearing early. Real fleets look different:

  • Few failures. Most assets never fail inside the observation window, so the positive class is tiny and concentrated in a few modes.
  • Right censoring. Assets are replaced, sold or decommissioned before failure, and a missing failure is not a negative label.
  • Interventions. Preventive maintenance, lubrication and alignment reset the degradation curve, and a model that has not seen them treats each reset as a mystery.
  • Label lag. The failure date in the CMMS is usually when the work order closed, not when the fault began.

Real labeled data does exist in the open literature. One 2025 dataset covers 93 district heating substations with fault labels validated against service reports, and it scores detectors on how early they flag a fault [2]. That design is the model to copy for industrial assets: maintenance records serve as ground truth, and evaluation rewards lead time, not just a classification hit.

Where real failure labels come from: the work order

The work order, not the sensor stream, is the label source, so buyers should evaluate label provenance before signal quality. In a CMMS or EAM system (SAP PM notifications and orders, IBM Maximo work orders, or similar), a corrective order typically carries the asset or functional location, a problem or failure code, a cause code, a remedy or activity code, free-text technician notes, and dates for report, start and completion. Those fields become the label, and their quality varies widely by site.

Ask four questions about every code set:

  1. Who assigned the code? A reliability engineer after root cause analysis, a technician at close-out, or a default value in a required field.
  2. When was it assigned? At notification (a symptom guess) or at completion (after the part was inspected).
  3. Is a taxonomy enforced? ISO 14224 defines equipment boundaries, failure modes, failure causes and maintenance actions so that reliability data can be compared across plants and owners. Sites that follow it, even loosely, need far less harmonization.
  4. Are detection and repair separated? You need the date the fault was detected, not just the repair date, to build RUL targets and lead-time metrics.

Free-text notes are often more accurate than the codes. "Replaced DE bearing, outer race spalling, found grease contamination" carries a failure mode, a cause and a component that a picklist of "Mechanical failure / Other" does not. Budget for text-to-code relabeling, and treat code noise as real: audits of widely used ML test sets estimate an average label error rate of at least 3.3%, enough to change model rankings [4]. Maintenance code sets rarely get that kind of audit, so measure their noise on a sample yourself.

How to specify assets, signals and sampling

A predictive maintenance data request should name the asset class, signal types and sampling regime precisely, because failure physics and sensor practice differ by equipment. Centrifugal pumps, induction motors, reciprocating and screw compressors, gearboxes and conveyors each fail in characteristic ways, and a model for one rarely transfers to another without retraining.

Signal specifications that change what a dataset is worth:

  • Vibration. Overall RMS velocity trends from a historian are cheap and common; raw waveforms or FFT spectra at kHz sampling are rarer and far more useful for bearing and gear faults. Ask for sensor location (drive end, non-drive end, axial), mounting and sample rate.
  • Process and electrical. Temperature, pressure, flow and motor current from PI, a SCADA historian or PLC tags, usually at seconds-to-minutes resolution and often compressed with deadband or swinging-door algorithms.
  • Oil and wear debris. Periodic lab reports with particle counts, viscosity and element ppm: sparse but strongly predictive for gearboxes and hydraulics.
  • Operating context. Run status, load, speed and setpoint, so the model can separate a fault from a change in duty.

For alarm and event streams that sit alongside these signals, see our guide to PLC, SCADA and DCS alarm and event logs. Downtime reason codes from production systems are covered in MES production and downtime records, and acoustic fault labeling has its own page on machine audio labeled with maintenance records.

Count labeled failure events, not rows

The value of a failure dataset is set by the number of labeled failure events per failure mode and asset class, not by row count or terabytes. A year of 1-second data from 500 pumps is billions of rows and may contain a dozen bearing failures. Buyers should ask suppliers for an event table before anything else: failure mode, asset class, count of events, count with usable pre-failure signal history, and count where detection date is known.

Three refinements matter:

  • Usable history window. An event counts only if the asset was instrumented and the historian was recording for the lead window you need (for example, 30 to 90 days before detection).
  • Censored units. Ask for assets that did not fail, with their in-service start and end, so survival and RUL models can use censored observations correctly.
  • Preventive and condition-based events. Include PM orders, inspections and component swaps with no failure found. Without them, a model can learn that a sudden drop in vibration after a bearing change is a precursor to failure, or that a planned replacement is a failure.

Avoiding leakage in train and test splits

Predictive maintenance data must be split by asset and by time, because random splits of overlapping windows leak the answer into the test set. Sliding windows from the same asset a few minutes apart are near-duplicates; if one lands in training and its neighbor in test, accuracy looks excellent and falls apart in production. Work on event-log benchmarks shows the same failure: random splits leak information across cases, and temporal, case-level splits are the common fix [3].

Practical rules for a held-out set:

  • Hold out whole assets, and ideally whole sites, so the model is tested on equipment it has never seen.
  • Cut training at a calendar date and test on later events only.
  • Leave a gap between the last training window and the first test window at least as long as your feature lookback.
  • Remove features derived from the work order itself (status, closure fields, technician notes) from model inputs, since they encode the label.

ISO/IEC 5259-4 frames this kind of discipline as part of a data quality process for training and evaluation data, including labeling for supervised learning [5]. It is a useful checklist when documenting how labels were derived for a customer or auditor. For evaluation designs that treat real business outcomes as ground truth, see outcome-labeled evaluation data.

Request specification for labeled failure data

A good request describes the data, the label rules and the minimum event counts, so a supplier can say quickly whether its records fit. Use the template below as a starting point and attach it to any sourcing conversation.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample specification
Asset classesCentrifugal pumps and induction motors, 50–500 hp
SignalsTriaxial vibration waveforms (min. 10 kHz) or spectra; bearing temperature; motor current; run status
Historian contextTag list, units, compression settings, timezone, gaps
Label sourceCorrective work orders with failure mode, cause and remedy codes plus technician text
Label datesNotification date, detection date if recorded, repair start, completion
Non-failure eventsPM orders, inspections, component replacements, decommissioning dates
TaxonomyNative code list plus mapping to ISO 14224-style modes, or permission to map
Minimum eventsAt least N labeled events per target failure mode with 60 days of prior signal
Delivery formatParquet for signals, CSV or JSON Lines for work orders, shared asset key
De-identificationTechnician names, badge IDs and site addresses removed or replaced

An illustrative joined label record might look like this:

{
  "asset_id": "PMP-0412",
  "asset_class": "centrifugal_pump",
  "work_order_id": "WO-88123",
  "order_type": "corrective",
  "failure_mode": "bearing_degradation",
  "cause_code": "lubrication_contamination",
  "remedy_code": "replace_component",
  "detected_at": "2025-03-04T06:10:00Z",
  "repair_completed_at": "2025-03-06T15:45:00Z",
  "signal_window_start": "2025-01-03T00:00:00Z",
  "label_assigned_by": "reliability_engineer",
  "technician_text_redacted": "DE bearing outer race spalling; grease contaminated"
}

The asset master also matters: install date, manufacturer and model, rated speed and power, and functional location. Work order histories on their own are covered on our maintenance work order datasets page and the maintenance logs page, and raw telemetry on the sensor and IoT data page. This page is about the join between them.

Rights, privacy and diligence questions for buyers

Equipment data looks impersonal, but work orders name people and plants, so rights and de-identification still need review before a license is signed. Technician names, badge numbers, supervisor sign-offs and free-text remarks about individuals are personal data in many jurisdictions. Plant locations, throughput figures and customer names inside notes can be commercially sensitive to the supplier.

Questions to settle with any supplier:

  • Does the operating company own the historian and CMMS data, or does an OEM, a monitoring service provider or a host site hold rights under a service contract?
  • Are signals from OEM-installed sensors covered by a separate data agreement?
  • Which fields are removed or pseudonymized, how, and was a sample checked?
  • What uses does the license allow: training, evaluation, benchmarking, inclusion in a product shipped to the vendor's customers?

SourceX sources operational datasets from US companies on request, rather than selling stock, and each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. If you are building failure prediction or RUL models and want to describe the data you need, you can submit a buyer request.

More context on manufacturing data sits on the industry-specific operational data hub and the manufacturing buyers page.

Sourcing labeled failure data for predictive maintenance

SourceX looks for US businesses that hold the condition-monitoring and work order data you describe, and every release is approved by the supplying company. The process runs from Find through Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees; a request does not guarantee a match. Describe your asset classes, signals and failure modes on the SourceX buyers page.

Frequently asked questions

Can simulated data like C-MAPSS still be useful?

Yes, for architecture prototyping and pretraining, but not as evidence a model will work in a plant. Simulated trajectories have controlled fault progression and no maintenance interventions, so validate on real labeled events before quoting accuracy to customers.

How should anomaly detection be evaluated on real data?

Score detections against work-order-confirmed faults and reward lead time before detection or repair, as the district heating dataset does [2]. Report false alarms per asset-month, since operators judge a system by alarm burden as much as recall.

Do I need ISO 14224 codes from the supplier?

No. Native codes plus technician text are enough if you can map them. ISO 14224 is a practical target taxonomy for harmonizing modes and equipment boundaries across sites.

Sources

  1. UCI Machine Learning Repository, "AI4I 2020 Predictive Maintenance Dataset" (2020). https://archive.ics.uci.edu/dataset/601
  2. arXiv, "Enabling Predictive Maintenance in District Heating Substations: A Labelled Dataset and Fault Detection Evaluation" (2025). https://arxiv.org/pdf/2511.14791
  3. arXiv, "Creating Unbiased Public Benchmark Datasets with Data Leakage Prevention for Predictive Process Monitoring" (2021). https://export.arxiv.org/abs/2107.01905
  4. arXiv (NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  5. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Data quality for analytics and machine learning - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data