Speech and audio data
Industrial Machine Sound Data for Acoustic Anomaly Detection
Quick answer
Machine sound data for anomaly detection should be real recordings from production assets: long runs of normal operation across every operating state, plus a smaller set of verified fault recordings for evaluation. Public benchmarks such as MIMII and the DCASE anomalous sound tasks are useful for method development, but they rely on a few machine types, staged faults and, in some cases, toy machines [1][2]. Specify machine taxonomy, microphone placement, sample rate, background noise and maintenance-linked labels before you license anything.
By SourceX Editorial · Updated
Why public machine sound benchmarks are not enough for production models
Public benchmarks prove that a method works on a curated test, not that it will hold up on your plant floor. The MIMII authors built their dataset because no public set covered industrial machine sound under normal and anomalous conditions in real factory environments, and even MIMII covers four machine types: valves, pumps, fans and slide rails [1]. Its anomalies, including contamination, leakage, rotating unbalance and rail damage, were introduced deliberately, not captured as they developed in service [1].
The DCASE 2020 Task 2 page states the core problem directly: anomalous sounds are rare and hard to collect, so systems are trained on normal sound only [2]. Part of its data comes from ToyADMOS, which uses miniature machines with deliberately damaged parts [2]. That design is right for a research challenge and wrong as the sole evidence base for a model that will page a reliability engineer at 3 a.m.
Three gaps usually appear when teams move from benchmark to deployment:
- Fault realism. Staged faults are often abrupt and severe. Real bearing wear, cavitation onset or belt slip starts subtle and changes over weeks.
- Asset diversity. Your fleet has specific makes, mounts, duty cycles and enclosures that a four-category benchmark does not represent.
- Acoustic environment. Neighbor machines, forklifts, HVAC, compressed-air bleed and PA announcements shape the noise floor more than any synthetic mix.
What to record: machines, operating states and microphone setup
A usable acoustic dataset is defined by the operating context of each clip, not just its audio. Capture machine type, model and serial (or a stable pseudonymous asset ID), operating state (speed, load, valve position, process step), and the run's start and end timestamps so you can join to historian or CMMS data later.
Microphone configuration decides what your model can learn. MIMII used a multichannel microphone array, which lets you test beamforming and source separation alongside single-channel detection [1]. For production data, ask for:
- Mic type and count: MEMS versus measurement-grade condenser, single versus array, and whether contact microphones or accelerometers were recorded in parallel.
- Placement: distance to the asset, mounting (magnetic, bracket, handheld), and whether placement is fixed across sessions. A mic moved 30 cm between sessions is a domain shift.
- Sample rate and bit depth: record at the native rate (often 44.1 or 48 kHz, 16- or 24-bit PCM) and downsample yourself. Many benchmark pipelines work at 16 kHz, which discards high-frequency content where some bearing and leak signatures live. Our audio file specs guide covers codec and channel trade-offs.
- Format: uncompressed WAV or FLAC. Reject MP3 or AAC for training; lossy codecs remove exactly the low-energy, high-frequency detail anomaly detectors depend on.
- Calibration: a reference tone or calibrator reading per session if you need absolute sound pressure level, not just relative features.
Background noise and SNR change detection performance
Factory noise level is a first-order variable, so every clip should carry an estimated signal-to-noise ratio or at least a noise-condition tag. MIMII mixed machine sound with real factory background noise at several SNR levels to show the effect [1]. A later method study on machine sound reports detection AUC falling as SNR drops from 6 dB to -6 dB [3].
In practice, ask suppliers to record a few minutes of ambient noise with the target machine stopped at each mic position and shift. That noise-only track lets you estimate SNR, build realistic augmentation, and separate "the machine changed" from "the room changed." It also surfaces intermittent sources, such as a neighboring press or a scheduled purge, that will otherwise show up as false positives.
Domain shift: plan for operating and environmental changes
Domain shift is a common reason a well-scored acoustic model fails after deployment. Public benchmarks were built to test whether methods can learn normal sound and flag deviations [2], but a model tuned on one load, one mic position and one room can treat an ordinary change in conditions as a fault. Research datasets released after the first DCASE tasks added deliberately shifted recording conditions for this reason.
For buyers this means the request should name the shifts you care about: speed or load changes, product changeovers, seasonal temperature, a relocated mic, a sister plant with the same equipment. Ask for each shift to be recorded as a separate, labeled condition rather than mixed into one pool. The same real-versus-simulated logic applies to room acoustics, as covered in our far-field speech data comparison.
Labels: normal-only training, verified anomalies for evaluation
Most acoustic anomaly detectors train on normal audio only, so the labels that matter most are the ones that certify "normal" and the ones that confirm a fault. The DCASE framing of training on normal data reflects this [2]. The same logic drives normal-only image sets for visual anomaly detection.
Ground truth for faults should come from maintenance evidence, not from someone listening to clips. Useful sources include:
- CMMS work orders (for example, SAP PM or Maximo notifications) with failure codes, component and close-out notes.
- Vibration route data or online vibration alarms on the same asset and time window.
- PLC and SCADA alarm history, which can mark trips, interlocks and process upsets. See PLC, SCADA and DCS alarm and event logs for how those records are structured.
- Teardown findings after a repair, which confirm the failure mode and let you label backward in time.
A clip recorded two weeks before a confirmed bearing replacement is a "pre-failure" candidate, not a certain anomaly. Keep three classes in your schema: verified normal, verified anomalous, and uncertain. Contaminated "normal" training data, where an early fault is already audible, is a silent failure mode that inflates false negatives.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"clip_id": "plantA-pump07-2026-03-14T02-10-00Z-ch3",
"asset_id": "PUMP-0007",
"machine_type": "centrifugal_pump",
"make_model": "redacted_by_supplier",
"operating_state": {"rpm": 1780, "load_pct": 72, "process_step": "transfer"},
"audio": {"format": "FLAC", "sample_rate_hz": 48000, "bit_depth": 24, "channels": 4, "duration_s": 10.0},
"mic": {"type": "MEMS_array", "position_id": "P2", "distance_m": 0.5, "mount": "magnetic_fixed"},
"noise": {"ambient_track_id": "plantA-P2-shiftB-ambient", "est_snr_db": 2.5},
"domain": "target_load_change",
"label": "uncertain",
"label_evidence": {"work_order": "WO-redacted", "failure_code": "BRG-WEAR", "days_to_repair": 13},
"speech_screened": true
}
Request template for acoustic condition monitoring data
A clear request lets a supplier say quickly whether they hold matching recordings. Fill in this template before talking to any source.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | What to specify | Why it matters |
|---|---|---|
| Machine types | Pumps, fans, compressors, gearboxes, valves, conveyors; makes if known | Fault signatures are type-specific |
| Operating states | Speed and load ranges, start-up and shutdown, changeovers | Normal variability must be covered |
| Normal volume | Hours per asset per state | Drives normal-only training |
| Fault evidence | Failure modes wanted and required proof (work order, teardown) | Separates real from staged faults |
| Mic setup | Type, count, placement, fixed or moved | Defines domain and array use |
| Audio spec | Native sample rate, bit depth, WAV or FLAC | Avoids lossy or resampled audio |
| Noise | Ambient tracks, SNR estimate, shift tags | Enables SNR-aware evaluation [3] |
| Domain splits | Named shifts recorded separately | Tests robustness to change |
| Side channels | Vibration, PLC tags, historian joins | Supports labels and multimodal models |
| Rights | Facility, equipment and speech checks | See next section |
If you also need vibration, temperature or current signatures, the sensor and IoT data page covers those streams, and manufacturing quality records can tie acoustic events to downstream defects. Industrial camera data for the same lines is covered in industrial defect image datasets.
Rights and confidentiality checks for factory audio
For machine sound, the main rights questions are usually facility confidentiality and equipment contracts, not personal data. Plant recordings can reveal throughput, process steps or line configuration, so confirm the supplying company has authority to license audio from that facility and that no customer or tolling agreement restricts it. Check whether OEM service contracts or remote-monitoring agreements give an equipment vendor rights over sensor data collected on its machines.
Incidental speech is the personal-data exception. Workers talk near machines, and radios and PA systems carry names. Ask for voice activity screening with removal or muting of speech segments, and see our BIPA and voiceprint risk guide for why voice content raises biometric concerns. If operators annotated recordings, the employee-authored data rights checks apply to their notes.
How SourceX handles machine sound requests
SourceX sources operational datasets from US companies and manages the commercial process, including licensing and ongoing purchases; recordings of hands-on work are among the data types it covers. Data is sourced on request, not held in stock, and a request does not guarantee a match. You describe the recordings you need, SourceX looks for US businesses that hold them, and every release is approved by the supplying company. More on the buying process is at SourceX for AI data buyers, and industry context is on manufacturing buyers.
Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows after an executed agreement and supplier approval. SourceX does not train models and does not publish prices.
Request real machine sound recordings for anomaly detection
If public benchmarks no longer reflect your assets, describe the machine types, operating states, recording setup and fault evidence you need. SourceX looks for US companies that hold matching operational recordings, reviews rights and agrees allowed uses in a license before anything is delivered. Start your request at SourceX for buyers, and browse related guides in the speech and audio data hub or the AI data guides.
Frequently asked questions
Can I train only on public datasets like MIMII and ToyADMOS?
You can develop and compare methods on them, but they cover few machine types and use staged or toy-machine faults [1][2]. Validate on recordings from your own asset types and acoustic environments before deployment.
How many fault recordings do I need?
Detectors trained on normal audio need anomalies mainly for evaluation and threshold setting. Prioritize coverage of failure modes and domains over raw fault count, and require maintenance evidence for each fault clip.
Should I buy audio or vibration data?
Often both. Microphones are non-contact and cheaper to deploy widely, while accelerometers are less sensitive to airborne background noise. Time-aligned audio and vibration from the same asset make labels easier to verify. This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Sources
- Purohit et al. (Hitachi), arXiv:1909.09347, "MIMII Dataset: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection" (2019). https://arxiv.org/pdf/1909.09347
- DCASE Community, "Unsupervised Detection of Anomalous Sounds for Machine Condition Monitoring (DCASE 2020 Task 2)" (2020). https://dcase.community/challenge2020/task-unsupervised-detection-of-anomalous-sounds
- arXiv:2412.10792, "Audio-based Anomaly Detection in Industrial Machines Using Deep One-Class Support Vector Data Description" (2024). https://arxiv.org/html/2412.10792v1
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.