Skip to content

Multimodal and embodied data

Time Synchronization in Multimodal Datasets: What Buyers Should Verify

Quick answer

Multimodal data synchronization is only as good as the clock behind each stream. Before accepting a delivery, buyers should know, per stream, whether timestamps come from a hardware trigger, a PTP or NTP disciplined clock, or software receive time; which epoch and timestamp field is authoritative; and how offset and drift were measured. Then test it yourself: detect shared events, cross-correlate signals, and compare the residual against a tolerance set for your application.

By SourceX Editorial · Updated

Why alignment fails even when every file has timestamps

Timestamps exist in almost every recording, but they rarely mean the same thing across streams. A camera may stamp a frame at exposure start, a driver may stamp it when the USB buffer arrives, and a logger may stamp it again when the message is written to disk. Each of those is a valid "time" with a different latency, and the gap between them is often tens of milliseconds and not constant.

Robotics containers make this explicit. In MCAP, each message carries both a log_time and a publish_time as nanosecond counts from an epoch the recording system chooses, which can be Unix time or robot boot time [1]. A delivery that mixes streams stamped against boot time with streams stamped against wall-clock time will look complete and still be unalignable without a documented offset. Our guide to robotics and sensor-fusion formats such as MCAP, ROS bags and RLDS covers the container side; this page covers the clocks.

Scale compounds the problem. Large egocentric and in-cabin collections mix many devices and sessions; Ego4D, for example, spans 3,670 hours from 931 camera wearers in 74 locations [5], and the DMD driver-monitoring set records RGB, depth and IR video from three cameras [4]. Every additional device is another clock to reconcile.

Hardware triggers, PTP, NTP and software timestamps compared

The synchronization method determines the best accuracy a dataset can ever reach, so record it per stream rather than per dataset. Methods fall into four tiers.

  • Hardware trigger or genlock. A shared pulse starts each camera exposure or ADC sample, so frames are captured at the same instant by construction. Residual error is usually trigger propagation and exposure-time differences. Ask for the trigger wiring diagram and which devices were free-running.
  • PTP (IEEE 1588). The Precision Time Protocol disciplines every device clock on a network to a grandmaster clock, and with hardware support it reaches far tighter agreement than NTP. Accuracy depends on hardware timestamping in NICs and switches; software-only PTP is much looser. Ask which device was grandmaster and whether it changed during a session.
  • NTP-disciplined clocks. Adequate for coarse alignment such as pairing logs with text, but error depends on network conditions and can step when the daemon corrects. Ask whether clocks were slewed or stepped during capture.
  • Software receive timestamps. The host stamps data on arrival. This includes driver, buffering and scheduling latency, which varies with CPU load, so offsets are neither zero nor constant.

A trustworthy supplier can say which tier each stream used and where the timestamp was taken in the pipeline. "Everything is synchronized" with no method named should be treated as software timestamps until proven otherwise.

How clock drift and offset show up in delivered data

Offset is a constant shift between two streams; drift is an offset that grows over time because two oscillators run at slightly different rates. A clock that is off by 20 parts per million accumulates 72 milliseconds of error per hour, which is invisible in a 30-second clip and obvious in a two-hour session. Long recordings therefore need drift measured at the start and end, not one offset.

Audio exposes drift through sample-rate mismatch. A recorder nominally at 48 kHz that actually runs slightly fast will slide against video frame by frame, and resampling to a nominal rate without correcting for the measured rate bakes the drift into the delivered file. Transcript timing has its own version of the problem: Whisper\u0027s utterance-level timestamps can be inaccurate and drift on long-form audio, which is why WhisperX adds forced phoneme alignment for word-level timing [2].

Other failure modes recur in acceptance testing:

  • Variable frame rate. Phone and screen recordings often use VFR; aligning by frame index times nominal fps assumes a constant interval and fails. Use per-frame presentation timestamps (PTS) instead.
  • Dropped or duplicated frames. A gap in a 30 fps stream shifts every later frame if alignment is index-based. Check inter-frame intervals against the nominal period.
  • Clock steps. An NTP correction mid-session produces a sudden jump, sometimes backwards, in otherwise monotonic timestamps.
  • Re-encoding. Transcoding can reset or regenerate PTS and drop the container\u0027s original timing metadata.
  • Timezone and epoch errors. Local time mixed with UTC, or boot-relative time mixed with Unix time, produces offsets of whole hours or decades.

Tolerances depend on the application, not the dataset

There is no single acceptable sync error; the tolerance comes from what the model will learn from the pairing. Speech-to-lip models and talking-head generation need frame-level audio-visual alignment, while robot control policies need sensor and action streams aligned within one control cycle. Pairing operator logs or incident notes with telemetry, as in paired time-series and text data, may tolerate seconds.

Write the tolerance into the request so the supplier can state whether their capture method can meet it. The multimodal dataset specification template has a place for it. ISO/IEC 5259-3 frames data quality as a managed process with requirements rather than prescribing metrics [6], so a stated tolerance and test method is how you turn "synchronized" into a measurable quality characteristic.

Illustrative example: invented to show structure; it does not describe an available dataset.

ApplicationStreams pairedExample tolerance to statePrimary test
Audio-visual speech, lip syncVideo, audioWithin one video frameClap or slate event, audio-visual offset estimate
Meeting understandingVideo, audio, transcript, screen shareWord timing within a few hundred msForced alignment of transcript to audio
Sensor fusion (camera, IMU, lidar)Image, IMU, point cloudSub-millisecond to a few msIMU vs. optical-flow motion correlation
Robot manipulation policiesCamera, joint states, actionsWithin one control cycleCommanded vs. measured joint motion lag
Log and text pairingTelemetry, operator notesSeconds to minutesKnown-event lookup in both streams

Acceptance tests to run on delivery

Run tests on a stratified sample of sessions across devices, sites and recording lengths before accepting the full delivery. The goal is to measure offset and drift independently of what the supplier reports, then compare.

  1. Monotonicity and gaps. For each stream, confirm timestamps never decrease and compute inter-sample intervals. Flag intervals beyond 1.5 times nominal as drops and near-zero intervals as duplicates.
  2. Epoch reconciliation. Confirm all streams declare the same epoch, or that a documented per-session offset exists. In MCAP, check which of log_time or publish_time the supplier treats as authoritative [1].
  3. Shared-event detection. Find events visible in several modalities: a clap or slate (audio spike plus visible contact), a flash (luminance step plus photodiode), or a robot contact (force spike plus visible touch). Measure the timestamp difference at the start and end of each session.
  4. Signal cross-correlation. Cross-correlate audio tracks from two microphones, or IMU angular velocity against camera optical-flow magnitude, over sliding windows. A peak lag that changes across the session is drift.
  5. Transcript alignment. Force-align a sample of transcripts with an aligner such as the Montreal Forced Aligner [3] and compare its word boundaries with the delivered timestamps.
  6. Drift model check. Fit offset against elapsed time per session. A non-zero slope means drift; a discontinuity means a clock step.

Record results per session in the acceptance report and tie them to the files you checked; our guide to dataset manifests and checksums explains how to pin a test result to exact bytes. For robot data specifically, combine these with the broader robot dataset acceptance checks.

Synchronization metadata to request with every delivery

The cheapest fix for alignment problems is metadata captured at recording time, because offsets measured later are estimates. Ask for a per-session sync manifest alongside the media; the SourceX sample manifest shows the general shape of delivery documentation.

Illustrative example: invented to show structure; it does not describe an available dataset.

session_id: s-000417
reference_clock: ptp_grandmaster   # stream all others align to
epoch: unix_utc_ns
streams:
  - id: cam_front
    device: global-shutter camera
    sync_method: hardware_trigger
    timestamp_point: exposure_start
    nominal_rate_hz: 30
    frame_rate_mode: constant
  - id: mic_array
    sync_method: ptp_hw_timestamp
    timestamp_point: buffer_first_sample
    nominal_rate_hz: 48000
    measured_rate_hz: 48000.6
  - id: imu
    sync_method: software_receive
    timestamp_point: host_arrival
    known_latency_ms: 4.2
measured_offsets_ms:
  - { stream: imu, vs: cam_front, start: 3.9, end: 4.6, method: imu_optical_flow_xcorr }
clock_events:
  - { t: \"2026-03-02T14:11:05Z\", type: ntp_step, delta_ms: -38 }
dropped_samples: { cam_front: 12, mic_array: 0, imu: 0 }

The fields that matter most are sync_method and timestamp_point per stream, the declared epoch, a measured rather than nominal sample rate for audio, and a log of clock events. If a supplier cannot produce these for historical recordings, ask how offsets would be estimated after the fact and on what sample.

How sync quality fits into contracts and governance

Synchronization belongs in the dataset specification and acceptance criteria, not only in an engineering conversation after delivery. State the tolerance, the test method and the sample size, and agree what happens to sessions that fail: re-alignment by the supplier, exclusion, or acceptance with a quality flag. Our guide to evaluating data supplier quality covers how to judge whether a supplier\u0027s own QA claims are credible.

For teams building high-risk AI systems under the EU AI Act, Article 10 requires training, validation and testing data to meet quality criteria and be subject to data governance and management practices [7]; as of October 2026, the high-risk application dates have reportedly been moved (to 2 December 2027 for Annex III and 2 August 2028 for Annex I) by Regulation (EU) 2026/1744 [8]. Documented sync methods and measured offsets are the kind of evidence that supports those practices for multi-stream data. See the training data quality hub for modality-agnostic checks.

Recordings that combine faces, voices and screens also raise privacy and rights questions that sync work can disturb: re-encoding to fix alignment should not undo redaction. Pair this page with de-identifying multimodal records and multimodal meeting recordings, or return to the multimodal data hub.

Where SourceX fits for synchronized recordings

SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and manages the licensing process; nothing is held in stock and a request does not guarantee a match. You describe the streams, sync method and tolerance you need, and every release is approved by the supplying company. You can describe your multimodal data requirement to SourceX at any stage of specification.

Request time-aligned multimodal data

Describe the modalities, synchronization method and alignment tolerance your model needs, and SourceX looks for US businesses that hold matching recordings. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Start a buyer request.

Sources

  1. Kaitai Struct (mirror of the MCAP specification), \"MCAP: format specification\". https://formats.kaitai.io/mcap
  2. Bain, Huh, Han, Zisserman (University of Oxford), arXiv:2303.00747, \"WhisperX: Time-Accurate Speech Transcription of Long-Form Audio\" (2023). https://arxiv.org/html/2303.00747v2
  3. McAuliffe et al., ISCA Archive (Interspeech 2017), \"Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi\" (2017). https://www.isca-archive.org/interspeech_2017/mcauliffe17_interspeech.html
  4. Ortega et al. (Vicomtech and collaborators), arXiv:2008.12085, \"DMD: A Large-Scale Multi-Modal Driver Monitoring Dataset for Attention and Alertness Analysis\" (2020). https://arxiv.org/abs/2008.12085
  5. Grauman et al. (Ego4D consortium), arXiv:2110.07058, \"Ego4D: Around the World in 3,000 Hours of Egocentric Video\" (2021). https://arxiv.org/abs/2110.07058v1
  6. ISO/IEC JTC 1/SC 42, \"ISO/IEC 5259-3:2024 Data quality for analytics and machine learning - Part 3: Data quality management requirements and guidelines\" (2024). https://www.iso.org/standard/81092.html
  7. European Commission, AI Act Service Desk, \"AI Act Article 10: Data and data governance\". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  8. Official Journal of the European Union, \"Regulation (EU) 2026/1744 of the European Parliament and of the Council\" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data