Video data
Frame Rate, Resolution and Codec Requirements for Video Training Data
Quick answer
Set video requirements from the task, not from a generic "1080p, 30 fps" default. Specify a minimum native frame rate tied to how fast the action moves, a minimum resolution tied to the smallest object or text your model must read, and a constant-frame-rate timebase. Ask for camera originals or a near-lossless master instead of a re-encoded copy, decide whether audio stays, and record HDR transfer functions. Write all of it as measurable acceptance criteria.
By SourceX Editorial · Updated
Frame rate: match it to the fastest motion you need to label
The minimum frame rate should be high enough that the shortest event you care about spans several frames at native capture. Many recognition models sample sparsely at training time, and dual-pathway designs such as SlowFast, which pair a low-rate pathway for scene content with a high-rate pathway for motion, only help if the source carries the temporal detail. You can always subsample 60 fps down to 15 fps; you cannot recover motion that was never captured.
Use the event duration to set the floor. A fastener click, a scanner trigger or a finger tap on a touchscreen can last a fraction of a second, so at 15 fps it may land in one or two frames, often motion-blurred. Slow processes such as a forklift approach, a lab incubation step or a patient transfer are well served at 15 to 30 fps.
Practical defaults that many CV teams start from (adjust after a pilot on your own model):
- Fine hand manipulation, assembly, surgical instruments: 30 fps minimum, 50 to 60 fps preferred, short shutter to limit blur.
- Whole-body actions, warehouse picking, sports-like motion: 25 to 30 fps.
- Slow processes, occupancy, step-level procedure tracking: 10 to 15 fps is often enough; 30 fps keeps options open.
- Vehicle and dashcam footage at speed: 30 fps minimum; rolling-shutter skew matters as much as rate.
Benchmarks such as Kinetics, built from short clips per class across 400 action classes, were collected from heterogeneous uploads rather than to a capture spec [1]. That is fine for pretraining breadth, but a purchase for a production model should state the rate explicitly. See temporal action segmentation labels for how frame rate interacts with boundary tolerance.
Variable frame rate breaks frame-indexed labels
Require constant frame rate (CFR) or, at minimum, preserved per-frame presentation timestamps, because variable frame rate (VFR) silently shifts labels. Phones, screen recorders and many body-worn or wearable cameras emit VFR by default: the container declares an average rate such as 30 fps, but actual frame intervals drift with load, light and thermal throttling.
The failure mode is quiet. An annotator labels "step 4 starts at frame 1,812" in a tool that decoded at the nominal rate, while your training loader seeks by timestamp, or vice versa, and the two disagree by hundreds of milliseconds late in a long clip. Audio-visual sync drifts the same way.
Ask suppliers to report, per file, r_frame_rate and avg_frame_rate from ffprobe, the stream time_base, and whether the two rates differ. If VFR is unavoidable, require labels in timestamps (seconds, or PTS in the stream timebase) rather than frame indices, and keep the original files so you can re-derive frames. Never accept a VFR-to-CFR conversion done after labeling unless the labels were remapped with it.
Resolution: size it from the smallest thing the model must see
Set resolution by the pixel footprint of the smallest object, gesture or text you must resolve, measured at the far edge of the scene, not by a display standard. A screw head, a gauge needle, a barcode or a 10-point UI label may need a minimum pixel height to be learnable, and that requirement drives capture resolution, lens choice and camera distance together.
Write it as a measurable rule: "the smallest labeled object is at least N pixels on its shortest side in 95% of labeled frames." That catches the common failure where 4K footage is delivered but the region of interest is a tiny corner of a wide-angle frame. Native resolution also matters: upscaled 720p labeled as 1080p is a frequent defect, detectable by inspecting high-frequency energy or comparing against a known camera model in the metadata.
Screen recordings are a special case. Text legibility needs native resolution (no scaling between the display and the capture), lossless or very lightly compressed encoding, and no chroma subsampling below 4:4:4 if color-coded UI elements matter, because 4:2:0 smears colored text edges. See software screencast video datasets for task and consent details. For still-image equivalents of these rules, see image resolution and compression requirements.
Codec, bitrate and why re-encoding costs you
Ask for the camera original or a near-lossless intermediate, because every lossy re-encode adds a new generation of artifacts that your model can learn as signal. Blocking, ringing and smeared fine texture accumulate across generations, and transcoding pipelines often also rescale, change frame rate or convert color without saying so.
Most enterprise video sits in H.264/AVC or H.265/HEVC inside MP4 or MOV, from cameras, VMS exports, phones or meeting platforms. Those originals are acceptable when the bitrate is adequate for the content. Problems come from platform-transcoded copies (social uploads, LMS exports, VMS "evidence export" presets) that were re-encoded for streaming at low bitrate.
Request these per file, and treat them as acceptance criteria:
- Codec, profile and level (for example H.264 High, HEVC Main 10), from ffprobe.
- Average bitrate and bits per pixel per frame, so a 1080p30 file at a few hundred kbps is flagged before you pay to label it.
- GOP structure and keyframe interval; long-GOP security exports can make frame-accurate seeking unreliable.
- Generation history: is this the camera original, a VMS export, or a platform transcode, and with what settings.
- Pixel format and chroma subsampling (yuv420p, yuv422p10le and so on).
Should you re-encode for training? Usually yes, once, on your side, into whatever your data loader prefers (frames as JPEG or PNG, or a decode-friendly codec), from the best master you can get. That decision belongs to your pipeline, not the supplier's export preset. Packaging, containers and clip indexes are covered in the delivery guide, and accepted file types in what file formats AI buyers accept.
HDR, color and bit depth
If any footage is HDR, require the supplier to declare the transfer function and color primaries, because unflagged HDR decoded as SDR looks washed out and shifts the color distribution your model learns. The ITU-R BT.2100 recommendation defines HDR television with two transfer functions, PQ and HLG, 10- or 12-bit depth, and the wide BT.2020 color gamut.
Ask for color_transfer (for example smpte2084 for PQ, arib-std-b67 for HLG, bt709 for SDR), color_primaries and color_space per file. Then decide one policy: keep HDR as 10-bit with metadata, or tone-map to SDR in your pipeline with a documented operator. Mixed, unlabeled HDR and SDR in one dataset is the failure to avoid. iPhone HDR capture is a common source of unexpected HLG content.
Audio, device mix and rights scope
Decide in the specification whether the audio track is kept, muted or removed, because audio changes both the data and its rights scope. Speech in a recording can be a confidential communication; California, for example, requires the consent of all parties to record one [3]. Voices can also carry biometric and personal information beyond what the picture shows. If you do not need audio, ask for it removed at source rather than muted, and say so in the license scope. See rights layers in a video clip and recording workers on video.
Device mix is a modeling decision. If your deployment camera is fixed (one VMS model, one dashcam SKU), match it closely and record camera make, model, lens and mounting. If deployment is open-ended, ask for deliberate diversity and record it per file. Ego4D shows what that looks like at scale: footage from 931 camera wearers in 74 locations across 9 countries [2].
Specification template for a video data request
Put requirements into a table the supplier can answer per file, and add a sample QC step before the full transfer.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Requirement | How it is verified |
|---|---|---|
| Native frame rate | 30 fps minimum; 60 fps for hand manipulation clips | ffprobe r_frame_rate on every file |
| Frame rate mode | CFR, or VFR with PTS preserved and labels in seconds | r_frame_rate equals avg_frame_rate; frame-interval histogram |
| Resolution | 1920x1080 native minimum; smallest labeled object at least 24 px | ffprobe; sample of 200 labeled frames measured |
| Upscaling | None | Camera model in metadata; spectral check on sample |
| Codec and generation | Camera original or ProRes/FFV1 intermediate; no platform transcodes | Generation history field; encoder tag |
| Bitrate | Minimum bits-per-pixel threshold set per codec after pilot | Computed from file size, duration, resolution |
| Chroma | 4:2:0 acceptable; 4:4:4 for screen recordings | pix_fmt |
| HDR | Declared; PQ/HLG kept as 10-bit with metadata | color_transfer, color_primaries |
| Audio | Removed at source | No audio stream in ffprobe output |
| Device metadata | Make, model, lens, mount position per file | Sidecar manifest |
| Timebase for labels | Labels reference PTS or seconds, with stream time_base recorded | Label file spot-check against decoded frames |
Run the spec on a pilot sample before scaling, and audit labels on that sample too; even widely used benchmarks carry measurable label error [4]. The document dataset requirements spec and how to write a data request for suppliers show the same request discipline in other modalities.
How SourceX handles video requests
SourceX sources operational datasets, including new recordings of hands-on work, from US companies on request; footage is not held in stock and a request does not guarantee a match. Buyers describe the data they need, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded, and delivery runs through private, access-controlled workflows after an executed agreement. You can describe a video data requirement to SourceX using the table above as your starting point. For the wider landscape, start at the video data hub or the AI data hub.
Specify your video training data with SourceX
SourceX sources operational video, including new recordings of hands-on work, from US companies on request, and manages licensing and ongoing purchases for AI teams wherever they are based. Nothing is contracted until a supplier agrees, and every dataset is delivered under a license defining records, uses, term and delivery. Share your video specification with SourceX.
Frequently asked questions
What frame rate is enough for action recognition?
For whole-body actions, 25 to 30 fps native capture is a common floor; for fine hand or instrument motion, request 50 to 60 fps so short events span several frames. Subsample in your loader rather than at capture.
Is 4K always better than 1080p for training?
No. Resolution matters only relative to the smallest object you must resolve and the bitrate available. Heavily compressed 4K can carry less usable detail than well-encoded 1080p, and it multiplies storage and decode cost.
Can a supplier convert VFR footage to CFR for me?
They can, but the conversion duplicates or drops frames and must happen before labeling, or labels must be remapped with it. Keep the originals and timestamps either way.
Sources
- arXiv (Kay, Carreira, Zisserman et al.), "The Kinetics Human Action Video Dataset" (2017). https://arxiv.org/abs/1705.06950v1
- arXiv (Grauman et al.), "Ego4D: Around the World in 3,000 Hours of Egocentric Video" (2021). https://arxiv.org/abs/2110.07058v1
- California Legislature, "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN§ionNum=632
- arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.