Skip to content

Schemas, packaging and delivery

Packaging Video Datasets: Containers, Sidecar Metadata and Clip Indexes

Quick answer

The safest video dataset format for machine learning is the supplier's original files, untouched, plus three machine-readable layers: a per-file technical sidecar extracted with ffprobe, a clip index that marks segments by start and end timestamps instead of cutting new files, and a manifest with checksums. Record each stream's timebase and frame-rate mode so variable frame rate footage does not silently misalign labels. Re-encoding, trimming and frame extraction belong in your pipeline, where you control them.

By SourceX Editorial · Updated

Why should video be delivered in its original container?

Original containers should be delivered as-is because re-encoding with the lossy codecs used for delivery degrades footage and every cut on a non-keyframe either re-encodes or drifts. A supplier who "normalizes" screen recordings to 30 fps H.264 for convenience destroys exactly the properties a model may need: crisp text on screen, original cadence of cursor movement, and true timing between events. Your curation pipeline will make its own choices about resolution, sampling and splitting, as large video training efforts such as Panda-70M and Cosmos do on source footage [3][4].

Containers such as MP4 (ISO base media) and Matroska are envelopes, not codecs. Matroska, now an IETF standard (RFC 9559), can carry many video, audio and subtitle streams in one file. That matters for deliveries: a single MKV or MP4 may hold a second audio track, a timed-text track or a rotation matrix that an extraction script drops if nobody listed it.

Ask for these container-level guarantees in the delivery spec:

  • No transcoding. Files are byte-identical to what the source system produced, verified by SHA-256 in the manifest.
  • All streams preserved. Secondary audio, subtitles and data tracks stay in place, or are listed as deliberately removed.
  • Container metadata reviewed, not blindly kept. Phone and camera files often carry creation_time, device model and GPS tags in container metadata; these may need stripping under the privacy method, which should be a remux (stream copy), not a re-encode.

If a supplier must transform files, for example to blur faces or redact screens, the transformed file becomes a new derivative. Deliver it with a pointer to the original's ID and a note of the encoder settings used, and see anonymizing video datasets for AI for the redaction side.

What belongs in a video technical sidecar?

A technical sidecar is a JSON file per video, generated by a standard probe, that lets you plan decoding without opening the video. ffprobe can emit stream and format information as JSON with -show_streams -show_format -print_format json, and it can print per-packet fields such as pts_time and duration_time. Ship the raw probe output and a short normalized summary, so your loader reads the summary and your auditors can check it against the raw output.

The fields that most often break pipelines when missing:

Field (ffprobe stream or format)Why it matters
codec_name, profile, pix_fmtDecoder support; 10-bit yuv420p10le fails on some GPU decoders
width, height, sample_aspect_ratioNon-square pixels distort if ignored
side data displaymatrix / rotationPhone video is often stored sideways and rotated on playback
time_baseThe unit every timestamp in the stream is counted in
r_frame_rate, avg_frame_rateDisagreement between them often signals variable frame rate; confirm from packet durations
nb_frames (or counted frames)Frame-index labels need a real count, not duration times fps
start_time, durationNon-zero start times shift every timestamp
color_primaries, color_transfer, color_rangeHDR or full-range footage looks wrong if treated as SDR limited range
audio sample_rate, channels, channel_layoutAlignment with transcripts and speech models

For computer-use and screen data, add capture context the probe cannot see: display resolution and scaling factor, capture tool, and whether the cursor was rendered into the frame. Those fields decide whether click coordinates in an action log map onto video pixels. Our guide to turning screen recordings into action-labeled trajectories covers the event-log side.

How does variable frame rate break video labels?

Variable frame rate breaks labels whenever annotations are stored as frame numbers but the frames are not evenly spaced in time. Screen recorders, video-conferencing exports and phone cameras commonly emit VFR: they drop or duplicate frames when nothing changes or when the device is under load. A label saying "frame 900" means 30.0 seconds only if every frame lasted exactly 1/30 s; in VFR footage it may land seconds away from the event.

Models also depend on timing. SlowFast-style architectures sample the same clip at different temporal rates [1], and that sampling assumes the loader knows when each decoded frame occurred. A VFR file decoded as if it were constant rate will stretch and compress motion unpredictably.

Practical rules for VFR deliveries:

  1. Store labels in time, not frames. Use presentation timestamps in seconds (or integer ticks plus the timebase), not frame indexes.
  2. Record the timebase per stream. A time_base of 1/90000 and a pts of 2,700,000 is 30.0 s; without the timebase the integer is meaningless.
  3. Flag the frame-rate mode. A frame_rate_mode field of constant or variable, derived by comparing packet durations from ffprobe -show_packets, tells your loader whether a fast path is safe.
  4. Never "fix" VFR at the source. Converting to constant rate duplicates or drops frames and re-encodes; do it in your pipeline if you need it, after the timestamps are captured.

Timestamps in sidecars and indexes should follow the conventions in timestamps and time zones in delivered datasets: media-relative offsets for positions in the file, and ISO 8601 with an offset for wall-clock capture times.

Why use a clip index instead of cutting clips?

A clip index records segments as start and end timestamps against the original file, which keeps originals intact and lets you re-cut at any sampling policy. Physically cutting clips with stream copy snaps cut points to the nearest keyframe, so a 4-second clip may become 6 seconds; cutting precisely forces a re-encode. Both outcomes are worse than a pointer.

Timestamped segments over untrimmed video are already the norm in research annotation: ActivityNet Captions marks events with start and end times inside long videos rather than shipping trimmed files [2]. The same pattern serves operations footage and screen sessions, where one 90-minute recording may contain dozens of tasks, interruptions and idle stretches.

A clip index is best shipped as JSON Lines, one segment per line, UTF-8 without a byte order mark [6]. That format streams, diffs cleanly between versions and loads into Parquet or Arrow with one call.

A reference folder structure and record layout

A predictable folder layout lets your loader find originals, sidecars, indexes and annotations by ID without guessing. The layout below separates immutable media from metadata that may be revised between dataset versions.

Illustrative example: invented to show structure; it does not describe an available dataset.

delivery_v1.2/
  MANIFEST.sha256              # every file, SHA-256, byte size
  dataset.croissant.json       # dataset-level metadata
  README.md                    # scope, known issues, changes vs v1.1
  media/
    vid_000417.mp4             # original, untouched
    vid_000418.mkv
  probe/
    vid_000417.ffprobe.json    # raw ffprobe output
  sidecars/
    vid_000417.meta.json       # normalized technical + capture fields
  index/
    clips.jsonl                # one segment per line
  annotations/
    actions.jsonl              # time-based labels keyed to video_id
  records/
    release_map.jsonl          # video_id -> release record ID, no personal data

Illustrative example: invented to show structure; it does not describe an available dataset.

{"clip_id": "vid_000417_c03", "video_id": "vid_000417", "stream_index": 0,
 "start_s": 312.480, "end_s": 341.112, "time_base": "1/90000",
 "start_pts": 28123200, "end_pts": 30700080,
 "frame_rate_mode": "variable", "segment_type": "task",
 "label": "reconcile_invoice", "redaction_applied": "screen_text_blur_v2",
 "release_record_id": "rel_7f3c", "split": "train"}

Note what the record does and does not hold. It carries a release record ID that links to the supplier's consent or release documentation, so counsel can trace any clip back to its permission, but it carries no names, emails or worker identifiers. Personal data stays out of every metadata layer, not just the pixels.

Checksums, Croissant and annotations alongside the video

Video deliveries need the same verification and description layers as any other dataset, plus a few video-specific ones. A SHA-256 manifest lets you confirm every multi-gigabyte file arrived intact before decoding; see dataset manifests and checksums. Croissant, a JSON-LD vocabulary built on schema.org, can describe the file resources and the record structure of clips.jsonl so standard tools can load it [5]; our page on Croissant metadata for licensed datasets covers what to request.

For training at scale you may later reshard into tar archives. Keep that a downstream step: shards built from originals plus the clip index are reproducible, while shards delivered by a supplier hide the originals. WebDataset tar shards explains when shards make sense as a delivery format.

Annotations should key to video_id and use the same time units as the clip index. Bounding boxes or keypoints defined per frame need the exact frame timestamp alongside the frame number, for the VFR reasons above.

Acceptance checklist for a video delivery

Run these checks on a sample before accepting the full delivery; each failure maps to a specific remediation request.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Every file in media/ matches its SHA-256 and byte size in the manifest.
  • A fresh ffprobe run matches the shipped probe/ JSON for codec, resolution, timebase and duration.
  • Each sidecar states frame_rate_mode; spot-check VFR flags by comparing r_frame_rate and avg_frame_rate.
  • Rotation, sample aspect ratio and color metadata decode correctly in your loader.
  • Every clip's start_s and end_s fall inside the video's duration, and start_pts divided by the timebase matches start_s.
  • No container tags contain GPS coordinates, device serials or personal names after the privacy method.
  • Every video_id has a release_record_id, and no metadata file contains personal data.
  • Annotation timestamps line up with visible events on a random sample of clips.

Where these packaging rules fit in a sourcing request

Packaging requirements are easiest to secure when they are written into the request and the license before collection or export begins. Operational video, including screen sessions, warehouse footage and egocentric recordings of skilled work, usually sits in systems that were never built to export ML-ready data, so stating "originals, ffprobe sidecars, timestamp clip index, release record IDs" up front saves a remediation cycle. The delivery hub covers the other formats you may need alongside video, and rights layers in a video clip covers what the release records need to cover.

SourceX sources operational datasets, including new recordings of hands-on work, from US companies on request; categories are not inventory, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded, and delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. You can describe the video you need, with your packaging spec attached, through SourceX for AI data buyers. Category pages for video recordings, screen recordings and egocentric video of skilled manual work describe what those requests typically involve.

Request licensed video data packaged for training

Describe the footage, the packaging spec and the uses you need, and SourceX looks for US businesses that hold that data, then assesses data and licensing permissions before anything is agreed. Nothing is contracted until a supplier agrees, and each dataset is delivered under a license defining records, uses, term and delivery. Start a request at sourcex.si/buyers.

Sources

  1. Feichtenhofer, Fan, Malik, He (arXiv:1812.03982), "SlowFast Networks for Video Recognition" (2018). https://arxiv.org/abs/1812.03982v2
  2. Krishna et al. (arXiv:1705.00754), "Dense-Captioning Events in Videos" (2017). https://ar5iv.labs.arxiv.org/html/1705.00754
  3. Chen et al. (arXiv:2402.19479), "Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers" (2024). https://arxiv.org/html/2402.19479v1
  4. NVIDIA (arXiv:2501.03575), "Cosmos World Foundation Model Platform for Physical AI" (2025). https://arxiv.org/html/2501.03575v2
  5. Akhtar et al., MLCommons Croissant working group (arXiv:2403.19546), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  6. jsonlines.org, "JSON Lines". https://jsonlines.org/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data