Speech and audio data
Word-Level Timestamps and Forced Alignment in Speech Datasets
Quick answer
Require timestamps at the granularity your model consumes: segment or utterance boundaries for ASR training chunks, word boundaries for subtitles, keyword search and redaction, and phone boundaries for TTS duration models. Ask how timings were produced (forced alignment against a verified transcript, ASR-native timestamps, or human correction), what reference they were scored against, and how drift was checked on long files. Then verify a sample yourself before accepting delivery, because timing errors are silent until training fails.
By SourceX Editorial · Updated
Which timestamp granularity should a speech dataset include?
The right granularity is the finest level any downstream use needs, because coarser timings can be derived from finer ones but not the reverse. Segment-level timings (start, end, speaker, text per utterance) are enough to cut long calls into the short windows most ASR trainers expect. Word-level timings are needed for subtitle generation, word-confidence filtering, targeted audio redaction and aligning transcripts with screen or event logs. Phone-level timings matter mainly for TTS duration and prosody modeling, and for phonetic research.
| Granularity | Typical fields | Main uses | Usual production method | Common failure mode |
|---|---|---|---|---|
| Segment / utterance | segment_id, start_s, end_s, speaker, text | ASR chunking, diarization training, LLM transcript text | Human segmentation, VAD plus ASR, or alignment | Boundaries cut mid-word; silence padding inconsistent |
| Word | word, start_s, end_s, confidence, aligned flag | Subtitles, redaction, keyword search, confidence filtering | Forced alignment to a verified transcript, or ASR word timestamps | Unalignable tokens (numbers, disfluencies) dropped or collapsed |
| Phone | phone, start_s, end_s, parent word | TTS duration and alignment, phonetics | Forced alignment with a pronunciation lexicon or G2P | Lexicon mismatch for names, accents and jargon |
For a broader view of what a full speech deliverable contains, start from the speech and audio data buyer's guide. Timestamp granularity is separate from what is written (see verbatim vs clean transcription standards) and from who spoke, which is a diarization question.
How is forced alignment produced, and why does the method matter?
Forced alignment takes audio plus a known transcript and finds where each word or phone occurs; its accuracy depends on the acoustic model, the pronunciation dictionary and how closely the transcript matches the audio. The Montreal Forced Aligner (MFA) is a widely used open-source example: a Kaldi-based GMM/HMM aligner with triphone acoustic models and speaker adaptation that can be trained on the data being aligned, with pretrained models, pronunciation dictionaries and grapheme-to-phoneme (G2P) tooling for out-of-vocabulary words [5]. Neural alternatives align with CTC or phoneme-recognition models instead of HMMs.
ASR-native timestamps are a different thing. Many end-to-end ASR models decode long audio in fixed windows, and the timestamps they emit are tied to that decoding, so they can be coarse at word level and can drift across long files. Open-source pipelines such as WhisperX respond by segmenting audio with voice activity detection and then running a forced-alignment pass with a phoneme recognition model to recover word timings [6].
Some curation pipelines go further and ensemble timestamps from several models to improve alignment reliability, as in a 2026 curated children's speech corpus [1]. For a buyer, the practical point is that "time-stamped transcript" can mean any of these methods, and vendor listings often advertise time-stamped transcripts without saying how they were made or checked [4]. Ask.
What should you ask a supplier about how timings were made?
The core questions are what the timings were aligned against, by which tool and version, and whether any human verified them. A forced alignment against a human-verified transcript behaves very differently from word timings emitted by the same ASR model that wrote the transcript, because the second case inherits the model's recognition errors at exactly the hard spots: crosstalk, accents, numbers and domain terms.
- Reference text: Was the transcript human-produced or corrected before alignment, and to which standard (verbatim with disfluencies, or clean)?
- Aligner: Which tool and version (for example MFA, a CTC or phoneme-recognizer aligner, a commercial ASR API), with which acoustic model and lexicon?
- Verification: What share of segments did a human check, how were boundaries corrected, and is a per-word "aligned" or "interpolated" flag present?
- Unalignable tokens: How are numerals, spelled-out letters, laughter, overlapping speech and inaudible tags handled? Were they dropped, given zero duration or interpolated?
- Long-file handling: Was the file aligned whole or in chunks, and how were chunk seams reconciled?
- Redaction: If personal details were removed, were timings recomputed after the audio was masked or replaced?
Treat any answer of "automatic, not reviewed" as a reason to budget your own verification, not as a disqualifier. Acceptance testing for the words themselves is covered in acceptance testing supplier transcripts.
How do you check alignment quality on a delivered sample?
Check alignment quality by measuring boundary error against a small hand-corrected reference and by running automated sanity checks across the full delivery. The reference set does not need to be large: a few dozen files, stratified by speaker, channel, noise level and file length, with word boundaries hand-corrected in a tool such as Praat or ELAN, will show systematic offsets quickly. Report the share of word boundaries within your tolerance and the median absolute offset, separately per stratum.
Automated checks catch structural problems the reference sample may miss:
- Word end times never precede start times, words within a segment are monotonic, and segment bounds contain their words.
- No timestamp exceeds the audio duration, which catches sample-rate mismatches (for example sample-index timings computed on a 16 kHz resample but converted to seconds at an 8 kHz original's rate).
- The distribution of word durations has no spike at zero or at a fixed value, which indicates interpolated or collapsed tokens.
- The offset between timings and audio energy onsets is stable from the start to the end of each file; a growing offset means drift.
- Timing coverage: the share of transcript tokens that carry real (not interpolated) timings.
Sample rate and channel layout affect these checks directly; see audio file specs for speech datasets.
Where does alignment usually fail in real-world recordings?
Alignment fails most often on long files, at redaction boundaries, in overlapping speech and on tokens the lexicon does not cover. Long recordings such as hour-long calls or meetings are where chunked ASR drifts; a fixed offset that grows across a file is the signature. Corpus builders routinely re-derive and re-segment source material for new uses: LibriTTS rebuilt LibriSpeech audio and text for TTS because LibriSpeech's 16 kHz audio suited ASR but was too low for high-quality synthesis [3].
Redaction creates a specific trap. If names, card numbers or account numbers are replaced by silence or tones after alignment, the words and their timings may no longer correspond to the audio, and if they are replaced before alignment, the aligner may stretch neighboring words across the gap. Ask that redaction tags carry their own start and end times; how audio redaction affects speech model training and redacting spoken PII from call recordings cover the details.
Conversational speech adds overlap, backchannels and interrupted words, which single-stream aligners handle poorly. Names, accents and domain jargon produce out-of-vocabulary words, so check whether G2P was used and whether the lexicon was extended for the domain.
How should timing labels relate to diarization and delivery formats?
Word and segment timings should use the same clock and the same file reference as any diarization labels, so one speaker turn in the diarization file contains exactly the words attributed to that speaker. RTTM, the common diarization exchange format, is plain text with one turn per line and fixed columns for start time, duration and speaker ID [2]; note that it stores duration, not end time, so conversions must add the two. Mismatched turn boundaries between RTTM and word timings are a frequent source of speaker-attributed ASR errors.
Agree the delivery format in the license schedule: seconds as decimals with a stated precision, a single time base per file, and timestamps relative to the start of the audio file, not wall-clock time. If you also need call start times in real-world time, deliver those as a separate ISO 8601 field; see timestamps and time zones in delivered datasets. Phone-level alignments for TTS are often delivered as Praat TextGrid files, while word timings for ASR are easier to validate as JSONL alongside per-file metadata (speech dataset metadata fields).
Specification template: time-aligned transcript deliverable
A written specification prevents the most common dispute, which is a buyer expecting verified word timings and receiving unreviewed ASR timestamps. Adapt the record below and the acceptance terms to your use case.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"file_id": "call_000184",
"audio": {"path": "audio/call_000184.flac", "sample_rate_hz": 16000, "channels": 2, "duration_s": 1843.52},
"time_base": "seconds_from_file_start",
"alignment": {"method": "forced_alignment", "tool": "MFA", "tool_version": "3.x", "reference_transcript": "human_verbatim", "human_verified_share": 0.15},
"segments": [
{"segment_id": "s0012", "channel": 1, "speaker": "agent", "start_s": 41.20, "end_s": 47.86,
"text": "thanks for holding, I can see the order from [REDACTED_NAME]",
"words": [
{"w": "thanks", "start_s": 41.20, "end_s": 41.55, "conf": 0.97, "aligned": true},
{"w": "for", "start_s": 41.55, "end_s": 41.70, "conf": 0.95, "aligned": true},
{"w": "[REDACTED_NAME]", "start_s": 46.90, "end_s": 47.86, "conf": null, "aligned": true, "redaction": "tone"}
]}
]
}
Illustrative acceptance terms to negotiate: a stated share of word boundaries within your tolerance on a buyer-selected sample, zero monotonicity violations, an interpolated-token rate below an agreed ceiling, no measurable drift between the first and last ten minutes of long files, and RTTM turns consistent with segment speakers. The numbers belong to your use case; subtitle work and TTS duration modeling need tighter tolerances than ASR chunking.
Sourcing time-aligned speech data with SourceX
SourceX sources operational datasets, including support and sales histories, from US companies on request, so you describe the recordings and timing granularity you need rather than picking from stock, and a request does not guarantee a match. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, with personal details removed or replaced and the method recorded. Describe your alignment requirements on the SourceX buyer page, or see how a call transcript is defined, then start a data request.
Frequently asked questions
Are ASR word timestamps good enough for training data?
Often for chunking ASR training audio, less often for TTS or redaction. Native ASR timestamps can be coarse or drift on long audio, which is why pipelines often add a forced-alignment pass, and some ensemble several models [1].
Can I re-align a licensed dataset myself?
Usually, if the license allows derived annotations and you have the transcripts. Running MFA or a phoneme-based aligner on delivered audio and transcripts is a common way to normalize timings across suppliers; confirm the license covers this.
Does real-world conversational speech align worse than read speech?
Generally yes, because of overlap, disfluency and noise. See spontaneous conversational speech data for why buyers still need it.
Sources
- arXiv, "CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling" (2026). https://arxiv.org/pdf/2607.03670
- arXiv, "The Multimodal Information based Speech Processing (MISP) 2022 Challenge: Audio-Visual Diarization and Recognition" (2023). https://arxiv.org/pdf/2303.06326
- arXiv (Zen et al.; Interspeech 2019), "LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech" (2019). https://arxiv.org/abs/1904.02882
- Unidata (vendor page), "Call Center Audio Dataset". https://unidata.pro/datasets/call-center-audio
- MFA, "Montreal Forced Aligner User Guide" (2026). https://montreal-forced-aligner.readthedocs.io/en/v3.3.9/user_guide/index.html
- arXiv (Bain et al.), "WhisperX: Time-Accurate Speech Transcription of Long-Form Audio" (2023). https://www.robots.ox.ac.uk/~vgg/publications/2023/Bain23/bain23.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.