Speech and audio data
Labeling Machine Audio with Maintenance Records: Building Real Fault Labels
Quick answer
Real fault labels for machine audio come from joining recordings to the maintenance history of the same asset: work orders, failure codes, inspection findings and repair dates in the plant's CMMS. You map each confirmed failure event to audio windows before and after it, assign a fault class from a controlled taxonomy, and mark uncertain windows instead of forcing a label. Because true faults are rare, use these labels for evaluation first, then for supervised training once each class has enough confirmed events.
By SourceX Editorial · Updated
Why maintenance records beat simulated faults for audio labels
Maintenance records give you faults that happened on production equipment under production conditions, which is the gap public benchmarks leave open. The DCASE 2020 anomalous sound detection task framed the problem as unsupervised precisely because anomalous machine sounds are hard to collect in quantity [1]. The MIMII dataset recorded valves, pumps, fans and slide rails with anomalies such as contamination, leakage, rotating unbalance and rail damage [2], which is useful for method development but does not tell you how your fleet sounds in the weeks before a bearing seizes.
Domain shift is the second reason. MIMII DUE showed that detection performance can drop sharply when operational or environmental conditions, such as machine speed or background noise, differ from those seen in training [3]. Audio paired with real work orders carries those conditions with it, provided the supplier also exports operating context. For the broader case on tabular failure labels, see labeled equipment failure data for predictive maintenance.
What a maintenance record can and cannot tell you
A work order tells you that someone acted on an asset and, if coded well, why; it does not tell you when the fault began. That distinction drives the whole labeling design. The fields that matter most in a CMMS export (IBM Maximo, SAP PM, Fiix, eMaint, UpKeep and similar systems) are:
- Asset identifier and hierarchy: functional location, equipment number, parent asset. Without a stable asset key you cannot join audio at all.
- Work order type: corrective, emergency, preventive, predictive, inspection. Only corrective and emergency orders, plus inspections that record a finding, are fault evidence.
- Failure codes: problem, cause and remedy codes (SAP PM catalog codes or Maximo failure hierarchies). ISO 14224 offers a reference taxonomy of failure modes, mechanisms and causes that many reliability teams map to [9].
- Timestamps: reported, scheduled, actual start, actual finish. Reported time is the closest proxy to fault detection; finish time marks the return to service.
- Free-text long description and technician notes: often the only place that says "bearing noise on drive end" or "impeller cavitation."
- Parts consumed: a replaced bearing, seal or belt is strong corroborating evidence for the fault class.
What records cannot tell you: onset time, severity progression, or faults that were never noticed. A preventive order that replaces a worn part without a finding is ambiguous. Treat absence of a work order as "no recorded fault," not as "healthy."
How to map work-order events to audio windows
The core method is to anchor each confirmed failure event on its reported timestamp and assign labels to time windows relative to that anchor. A practical scheme uses four zones per event:
- Pre-failure window (for example, days to weeks before the reported time): candidate degradation audio, labeled with the fault class and a time-to-event value for remaining-useful-life models.
- Event window (reported time to actual start of repair): the fault is known to be present.
- Repair gap (actual start to actual finish): exclude; the machine is stopped, disassembled or running under test.
- Post-repair window (after finish, once the asset is back on its normal duty cycle): a within-asset "healthy" reference, which is the cleanest negative you will get.
Window lengths are a modeling hypothesis, not a fact of the data, so record them as parameters and test sensitivity. Bearing defects may be audible long before a failure is reported, while a sudden seal leak may not. Keep a buffer zone around the pre-failure boundary labeled "uncertain" rather than "normal," because early-stage faults in that zone would otherwise poison your negative class.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value | Purpose |
|---|---|---|
| clip_id | PMP-07_2026-03-14T02:10:00Z_10s | Unique audio window key |
| asset_id | PMP-07 | Join key to CMMS functional location |
| asset_class / model | centrifugal pump / OEM model string | Group-level splits and taxonomy |
| sensor_id / position | MIC-2 / drive-end bearing housing, 15 cm | Controls for microphone placement |
| clip_start_utc | 2026-03-14T02:10:00Z | Normalized clock, see alignment section |
| operating_state | running, 1,480 rpm, 72% load | Explains domain shift [3] |
| work_order_id | WO-118842 | Provenance of the label |
| wo_type / failure_code | corrective / BRG-NOISE, cause WEAR | Fault class source |
| label_zone | pre_failure | Pre, event, repair gap, post, uncertain |
| hours_to_event | 46.5 | Remaining-useful-life target |
| label_confidence | confirmed_by_parts | Confirmed by parts, by note, code only |
Building a fault taxonomy that survives contact with real codes
Start from the failure codes the supplier actually uses, then collapse them into classes your model can learn, rather than imposing a benchmark taxonomy top-down. Benchmark categories such as leakage, unbalance and contamination [2] are a sensible target vocabulary, but plant codes are messier: "NOISY," "VIBRATION HIGH," "TRIPPED" and "OTHER" often dominate.
A workable approach is a two-level mapping. Level one is the acoustic class (bearing defect, imbalance or misalignment, cavitation, leakage, belt or coupling wear, electrical hum, no fault found). Level two keeps the original code and free text so you can remap later. Assign "no fault found" orders to their own class; they are valuable hard negatives where a technician heard something and found nothing.
Expect the long tail. Two or three classes will hold most confirmed events, and many classes will have a handful. Report per-class counts before any modeling decision and merge or hold out classes that cannot support a stable evaluation.
Clock alignment, asset joins and other failure points
The most common way these datasets go wrong is a silent join error, and clock drift is the usual culprit. Edge recorders, PLC historians and CMMS servers keep separate clocks, often in different time zones, and work-order timestamps are frequently entered by hand at shift end. A two-hour offset can move audio from the event window into the post-repair window.
Checks to run before trusting any label:
- Time zone and DST normalization: convert everything to UTC and confirm how the CMMS stores local time.
- Run-state cross-check: if a historian or SCADA tag records motor run status, verify the repair gap shows the machine stopped. Real-world operational datasets such as CARE to Compare pair sensor data with turbine status information for exactly this kind of context [4].
- Asset relocation and swaps: rotating spares get moved between positions; the functional location and the serial-numbered equipment are not always the same thing.
- Sensor changes: a microphone moved, replaced or re-gained changes the recording more than many faults do. Ask for a sensor change log.
- Backfilled work orders: orders created days after the fact with a default timestamp. Flag orders whose reported and finish times are identical or fall outside shift hours.
Evaluation design: class imbalance, leakage and label noise
Use maintenance-derived labels first as an evaluation set, because they are scarce, noisy and expensive, and evaluation is where real faults matter most. The DCASE framing, training on normal audio and scoring anomalies [1], fits this well: you train on abundant normal data and measure detection against confirmed events. Report AUC and partial AUC per asset class [1], plus detection lead time in hours before the reported event.
Split by asset and by time, never by random clip. Adjacent clips from the same failure are near duplicates, and research on event-log benchmarks shows that random splits leak information across cases [8]. Hold out entire assets, or all events after a cutoff date, so the test set measures generalization to new failures.
Budget for label noise. An audit of widely used test sets, including audio, estimated an average label error rate of at least 3.3% even in curated benchmarks [5]; code-only maintenance labels will be noisier. Have a reliability engineer review a sample of each class and record agreement. ISO/IEC 5259-4 is a useful frame for documenting that labeling process for training and evaluation data [6]. The golden evaluation dataset from business records guide covers adjudication workflows in more depth.
What to request from a supplier
Ask for the raw evidence behind each label, not just a labels column, so you can rebuild labels under your own windowing rules. A complete request covers:
Illustrative example: invented to show structure; it does not describe an available dataset.
- Audio: format (WAV or FLAC), sample rate, bit depth, channel count and microphone model per sensor; see audio file specs
- Coverage: assets per class, recording hours per asset, and continuous versus scheduled snapshot capture
- CMMS export: work order ID, type, asset ID, failure problem, cause and remedy codes, reported, start and finish timestamps, long description, parts consumed
- Asset register: model, manufacturer, install date, functional location history
- Operating context: run status, speed, load or flow from the historian, aligned to UTC
- Sensor metadata: position, mounting, change log
- Clock documentation: time sources, time zones and any known offsets
- De-identification note: technician names and badge numbers removed from free text
- Metadata file describing schema and record structure, for example in Croissant JSON-LD [7]
Buyers of maintenance work order datasets and licensed maintenance logs should ask the same supplier whether condition-monitoring audio or sensor and IoT data exists for the same assets; the join is where the value is, and you can describe that paired request to SourceX's buyer team. For how this fits other audio sourcing decisions, start at the speech and audio data hub, and for outcome-anchored labels in general see outcome-labeled evaluation data.
Sourcing machine audio paired with maintenance records
SourceX sources operational datasets from US companies on request, including engineering records and new recordings of hands-on work, and looks for businesses that hold the data you describe. Nothing is in stock and a request does not guarantee a match; every release is approved by the supplying company and personal details are removed before delivery. Describe the assets, fault classes and record fields you need at SourceX for buyers.
Frequently asked questions
Can I use preventive maintenance orders as negative labels?
Only with care. A preventive order with no recorded finding suggests the asset was serviced, not that it was healthy beforehand. Treat the period after a clean preventive order as a weaker negative than the post-repair window of a corrective order.
How many confirmed fault events do I need?
There is no universal number. Count events per class per asset type, and if a class cannot fill a held-out test split with several independent events from different assets, keep it in evaluation only or merge it.
Is ambient factory audio enough, or do I need contact sensors?
Airborne microphones capture leakage, cavitation and belt noise well but pick up neighboring machines. Ask suppliers which capture method was used per asset and keep it as a feature, because mixing methods without recording them creates a hidden domain shift [3].
Sources
- DCASE Community, "DCASE 2020 Task 2: Unsupervised Detection of Anomalous Sounds for Machine Condition Monitoring" (2020). https://dcase.community/challenge2020/task-unsupervised-detection-of-anomalous-sounds
- arXiv (Purohit et al.), "MIMII Dataset: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection" (2019). https://arxiv.org/pdf/1909.09347
- arXiv (Tanabe et al.), "MIMII DUE: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection with Domain Shifts" (2021). https://arxiv.org/pdf/2105.02702
- arXiv, "CARE to Compare: A real-world dataset for anomaly detection in wind turbine data" (2024). https://arxiv.org/pdf/2404.10320
- arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
- arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- arXiv (Weytjens and De Weerdt), "Creating Unbiased Public Benchmark Datasets with Data Leakage Prevention for Predictive Process Monitoring" (2021). https://export.arxiv.org/abs/2107.01905
- ISO, "ISO 14224:2016 - Petroleum, petrochemical and natural gas industries — Collection and exchange of reliability and maintenance data for equipment" (2016). https://www.iso.org/standard/64076.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.