Multimodal and embodied data
Paired Time-Series and Text Data: Telemetry with Incident Notes, Alarms and Operator Logs
Quick answer
A time series and text dataset for model training is real sensor, process or service telemetry joined in time to the words people wrote about it: alarm messages, shift handover notes, maintenance work orders, incident tickets and postmortems. Its value is the pairing, not either modality alone. Buyers should specify the signal windows, the text sources, the clock rules that align them, and whether each note explains, observes or merely coincides with the signal, then verify those rules on a sample before licensing.
By SourceX Editorial · Updated
Why public corpora rarely contain real telemetry-text pairs
Most open time-series training data is numeric only, so text-grounded reasoning has to come from operational records. Google's TimesFM team describes a pretraining corpus built mainly from public numeric archives plus synthetic series [1], and benchmarks such as GIFT-Eval [3] and fev-bench [2] score forecasting accuracy rather than explanation. Neither tells a model why a compressor tripped.
Public text-plus-series collections that do exist tend to pair series with news or public reports, which raises two problems: the text is often only loosely related to the series, and it can contain forecasts or outcomes that leak the answer. The fix is to give every numeric and text record its own start and end time and to keep factual text separate from predictive text. For industrial copilots and root-cause models, the text you need sits inside historians, CMMS and ITSM systems at operating companies.
Which text sources pair well with which signals
The best pairs come from systems where a human wrote text because of something in the signal. Each source has a different timestamp meaning and a different noise profile:
- Alarm and event journals (DCS, SCADA, NMS). Short, templated messages with tag IDs, priority and state. ISA-18.2 state models distinguish unacknowledged, acknowledged, shelved and suppressed alarms [6], so an alarm row is a state transition, not a single event. Ask for the transition history, not just activations. Network equivalents are covered in NOC alarm and incident logs for root-cause AI.
- Shift handover logs. Free text written at shift change summarizing abnormal readings, open work and equipment out of service. The UK HSE treats handover as safety-critical communication [7], which makes these notes dense with causal language, but each one refers back to an 8 or 12 hour window.
- Maintenance work orders. Records from SAP PM, IBM Maximo or similar CMMS with failure codes, problem and cause descriptions and labor text. Request-created, actual-start and completion timestamps differ, often by days. See maintenance work order datasets and maintenance logs.
- Incident tickets and postmortems. In ServiceNow, work notes and customer comments live as separate journal entries in sys_journal_field, each with its own creation time [5], so a ticket is a time series of text. Postmortems add a reconstructed timeline and a stated root cause; see incident postmortems.
- Operator chat and paging threads. Highest temporal resolution, lowest structure, and the highest density of names and personal details.
How to define a pair: timestamp, window and relation
A usable pair needs three explicit fields: when the text was written, which signal window it refers to, and what relation it claims. Ambiguity on any of these makes training targets unreliable.
Authored time versus referenced time. A handover note written at 06:58 may describe a pressure excursion at 02:10. Ask suppliers whether the referenced window is recorded in a field, parsed from the text, or inferred by proximity. Proximity pairing (nearest signal within N minutes) is cheap and often wrong for handovers and work orders.
Relation type. Label each pair as one of: causal explanation (the note states a cause), observation (the note describes the signal without explaining it), action (the note records an intervention that changes the signal afterwards), or coincident (same time, no stated link). Models trained on undifferentiated pairs learn to narrate signals rather than diagnose them.
Information leakage. For root-cause or forecasting tasks, text written after an outcome must not be visible when the model predicts that outcome. The point-in-time join pattern from feature stores applies directly: Feast treats each row's event_timestamp as an inclusive upper bound, so only values known at that moment are returned [4]. Apply the same rule to text, using authored time, not referenced time. The companion page on telemetry aligned with incident labels for root-cause models covers label design in more depth.
Clock and time-zone normalization is the main failure point
Misaligned clocks between the historian and the ticketing system break more pairs than any other single issue. Historians such as OSIsoft PI (AVEVA PI) commonly store UTC with source-device timestamps; CMMS and ITSM systems often store server-local time or the user's display time zone; handover logs may be typed with wall-clock times and no zone at all.
Ask suppliers to document, per source system:
- the stored time zone and whether daylight-saving transitions are represented (a duplicated 01:00 to 02:00 hour in November is a common source of silent misalignment);
- whether timestamps are device time, gateway time or server ingest time, and the typical lag between them;
- known clock drift or NTP outages, and how they were corrected;
- the resolution of each timestamp (seconds, minutes, or date only for some work-order fields).
A quick verification test: pick 20 alarm activations with an unambiguous signal signature, such as a trip, and measure the offset between the alarm row and the signal edge. A consistent nonzero offset suggests a zone or ingest-lag error; a scattered one suggests drift. The general method is in time synchronization in multimodal datasets.
Illustrative pair record and specification
A pair record should carry both modalities' identifiers, both time semantics and the relation label, so buyers can rebuild windows and filter by relation.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"pair_id": "p-000184",
"asset_id": "ASSET_7F3A",
"site_id": "SITE_02",
"signal": {
"tags": ["ASSET_7F3A.DISCH_PRESS", "ASSET_7F3A.MOTOR_AMPS", "ASSET_7F3A.VIB_X"],
"window_start_utc": "2025-11-03T01:40:00Z",
"window_end_utc": "2025-11-03T03:10:00Z",
"sample_rate_s": 1,
"units": ["kPa", "A", "mm/s"],
"quality_flags": "per-sample OPC quality code retained"
},
"text": {
"source_system": "shift_handover_log",
"authored_utc": "2025-11-03T06:58:00Z",
"referenced_window_method": "parsed_from_text",
"author_role": "OPERATOR_ROLE_B",
"body": "Disch press dropped ~02:10, amps spiked, unit tripped on high vib. Reset 02:45 after checking ASSET_7F3A strainer. WO raised."
},
"relation": "observation+action",
"linked_records": ["WO_REF_5521", "ALM_SEQ_88310-88316"],
"leakage_cutoff_utc": "2025-11-03T06:58:00Z"
}
Illustrative example: invented to show structure; it does not describe an available dataset.
| Specification item | What to ask for | Red flag in a sample |
|---|---|---|
| Signal granularity | Native sample rate and any compression (PI exception or swinging-door) | Uniform 1-minute averages labeled as raw |
| Text sources | Named systems and tables (alarm journal, sys_journal_field, CMMS long text) | Only final ticket summaries, no journal history |
| Time semantics | Authored, referenced and ingest time per text record | One timestamp field with no stated meaning |
| Relation labels | Causal, observation, action, coincident; who labeled and how | All pairs created by proximity only |
| Alarm states | Full ISA-18.2-style transition history [6] | Activations only, no shelve or suppress rows |
| Negative windows | Normal-operation windows with no text, at known ratio | Every window contains an incident |
| Identifier mapping | Consistent pseudonyms across signal tags and text | Asset ID masked in text but clear in tag names |
Pseudonymization must be consistent across both modalities
Asset IDs, site names and people's names appear in tag paths and in free text, and the replacement must be identical in both or the pairing breaks. A tag such as "PLANT2.K-101.DISCH_PRESS" and a note reading "K-101 tripped again" need the same pseudonym, applied by a shared lookup before delivery, not by separate tools per modality.
Free text adds indirect identifiers that tag pseudonymization misses: a rare event on a named night shift, a small site, a job title held by one person. These quasi-identifiers are covered in indirect identifiers in business text. Technician names in work orders and paging threads are the most common direct identifier; ask how they were replaced and whether a sample was checked.
Supplier-side confidentiality matters too. Asset identifiers, process setpoints and site names can reveal a company's operations, so expect the supplying company to decide what is released and at what level of masking.
Splits, evaluation sets and contamination
Split by asset and by time, never by random pair. Random splits put the same incident's alarm, handover note and work order on both sides, and the model learns to recognize the event rather than reason about it. GIFT-Eval's attention to overlap between pretraining and test data [3] is the same concern at corpus level.
For evaluation, hold out whole sites or later periods, and include coincident and negative pairs so a model is scored on saying "the note does not explain this signal." Time-series question answering sets can be built from these holdouts by turning postmortem root-cause statements into questions over the referenced window. Keep such sets private if they will be used repeatedly; see private evaluation sets for multimodal models.
Rights and documentation for records assembled from several systems
A telemetry-text pair often combines data from different systems, vendors and sometimes contractors, so confirm who controls each part. Historian data may sit under an OEM service agreement; tickets may include text written by a managed-service provider; chat may fall under employee monitoring policies. The issues are covered in licensing multimodal records from several rightsholders.
If you want a supplier search run against a written specification like this one, you can send the requirements to SourceX. Either way, ask for documentation per dataset: source systems, extraction queries, time normalization steps, pairing method, relation labeling and masking method. Data Cards provide a workable structure for this [8]. Do not rely on a hosting page's license field for anything you plan to train on; an audit of popular hosting sites found license omissions and errors to be common [9].
Pure sensor licensing without text is covered in sensor and IoT data, and synchronized video plus machine signals in synchronized manufacturing data. For the wider landscape, start from the multimodal data hub or the structured and time-series data guide.
How SourceX can help with telemetry and text pairs
SourceX sources operational datasets from US companies on request, including engineering records and support histories, and manages licensing; nothing is held in stock and a request does not guarantee a match. Each dataset is rights-reviewed, personal details are removed or replaced before delivery with the method recorded, and every release is approved by the supplying company. To describe the signals, text sources and pairing rules you need, submit a buyer request at SourceX.
Sources
- Google Research, "A decoder-only foundation model for time-series forecasting" (2024). https://research.google/blog/a-decoder-only-foundation-model-for-time-series-forecasting/
- arXiv, "fev-bench: A Realistic Benchmark for Time Series Forecasting" (2025). https://arxiv.org/html/2509.26468v4
- arXiv, "GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation" (2024). https://arxiv.org/abs/2410.10393
- Feast, "Feature retrieval". https://docs.feast.dev/v0.63-branch/getting-started/concepts/feature-retrieval
- ServiceNow, "Journal fields". https://www.servicenow.com/community/developer-forum/how-to-delete-comment-or-work-notes-from-an-incident-task-using/m-p/3146067
- ICONICS, "ISA 18.2 Alarm State Transition Diagram". https://documentation.iconics.com/v10.97.3/Content/Alarming/Alarm Server/Alarm References/alarm-state-transition-diagram.htm
- UK Health and Safety Executive, "Shift handover". https://www.hse.gov.uk/humanfactors/topics/shift-handover.htm
- Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.