Video data
How Many Hours of Video Do You Need? Sizing a Video Dataset
Quick answer
There is no fixed hour count. Size a video purchase from three numbers, not raw footage: usable hours left after filtering and anonymization, the number of distinct scenes, sites and subjects, and labeled instances per class, especially the rarest class you must recognize. Public benchmarks set a reference point: Kinetics has at least 400 clips per action class, each from a separate video [1]. Then run a pilot, plot a learning curve, and buy the volume the curve justifies.
By SourceX Editorial · Updated
Why raw hours are the wrong unit for video
Raw hours overstate what a model learns because video is highly redundant from frame to frame and from day to day. Fixed-camera deployments, such as the urban intersection testbed described in [4], record one view continuously, so a camera filming the same intersection, assembly cell or loading dock for a month produces hundreds of hours that share one background, one lighting pattern and one camera angle. A model trained on that footage learns the scene, not the action, and fails when the camera moves.
Diversity is the real denominator. Kinetics deliberately drew each clip from a different YouTube video [1], and Ego4D reports its 3,670 hours alongside 931 unique camera wearers and 74 locations [2]. When a supplier quotes hours, ask for the same breakdown: how many distinct sites, cameras, subjects, shifts and recording days sit behind the total. The video data hub covers where real-world footage comes from.
The four counts that actually size a video dataset
Specify a video purchase with four counts, then derive hours from them rather than the reverse.
- Unique scenes. A scene is a camera position at a site. Ten cameras across ten facilities beat one hundred hours from one camera.
- Unique subjects. Count distinct people performing the actions, because models overfit to body shape, clothing and handedness when a few workers dominate.
- Per-class instances. Count labeled occurrences of each action, with start and end timestamps. Rare classes, such as a dropped part or an unsafe lift, set the floor for total volume.
- Usable hours. Count hours that survive quality checks, consent and anonymization review, and deduplication.
The per-class count usually dominates. If an unsafe-reach event occurs once per four hours of footage and you need several hundred instances for training plus a held-out set, you need well over a thousand raw hours of that scene type, even if common actions saturate in fifty. Kinetics' authors also tested whether class imbalance biased their classifiers [1], which is the same question you should ask of your own long-tail classes.
Delivered hours versus usable hours
Usable hours are what remain after you remove footage your pipeline cannot train on, and the gap from delivered hours can be large. Budget for the loss before signing, and agree with the supplier which side absorbs it.
Typical rejection causes in operational video include:
- Dead time: idle cameras, empty rooms, breaks and night shifts with no activity.
- Technical defects: dropped frames, variable frame rate, corrupt containers, timestamp drift between cameras, and resolution below your model's input size after cropping.
- Occlusion and framing: the action happens outside the field of view or behind equipment.
- Anonymization damage: blurring faces and license plates, as fixed-camera intersection datasets do [4], protects people but can erase cues your model needs, such as gaze or hand-face contact [6]. Check redaction on a sample before you commit.
- Rights and consent exclusions: segments showing people who did not consent, screens with third-party content, or recorded audio that a state law treats differently from video. See rights layers in a video clip and recording workers on video.
- Weak text alignment: for video-language work, narration or transcripts often describe something other than what is on screen. HowTo100M's captions come from narration, often ASR output, and are not manually annotated [5].
Sizing worksheet for a video data request
Use this worksheet to turn a model goal into a defensible volume number before you talk to suppliers.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Line | Field | Example value | How to get it |
|---|---|---|---|
| A | Target task | Temporal action localization, 14 assembly steps plus 3 error classes | Model spec |
| B | Rarest required class | "Fastener skipped," about 1 per 6 working hours | Supplier sample or site logs |
| C | Instances needed per class (train) | 300 | Pilot learning curve |
| D | Instances held out (eval) | 100 per class, from sites not in training | Eval plan |
| E | Raw hours to hit rarest class | (300 + 100) x 6 = 2,400 | C, D and B |
| F | Expected usable ratio | 0.55 (dead time, occlusion, consent exclusions) | Pilot rejection log |
| G | Delivered hours to request | 2,400 / 0.55, about 4,360 | E / F |
| H | Minimum unique sites | 12 (8 train, 4 held out) | Generalization target |
| I | Minimum unique subjects | 60, no subject over 5% of clips | Overfitting check |
| J | Clip spec | 1080p, 30 fps constant, H.264 MP4, UTC timestamps | Pipeline input |
Two lessons come out of a worksheet like this. The rarest class, not the model architecture, drives hours, and the held-out set must come from sites and subjects the model never saw in training, or your evaluation measures memorization. For how large that held-out set needs to be to detect a real difference, see sizing an eval set for statistical power.
Run a pilot and read the learning curve
A pilot with a learning curve is the only reliable way to set volume, because the required size depends on your model, labels and domain. Across domains, generalization error falls as a power law of training-set size [3], so a few points on a log-log plot let you estimate how much more data buys how much improvement.
A practical pilot sequence:
- License a small, diverse slice: many scenes and subjects, few hours each. The data pilot guide covers terms and logistics.
- Train on nested subsets at 10%, 25%, 50% and 100% of the slice, keeping the held-out sites fixed.
- Plot per-class metrics, not just the average. A flat average can hide a rare class that is still climbing steeply.
- Fit a line to the log-log curve and extrapolate to your target metric. If the curve is flattening, more of the same footage will not help; buy new scenes, subjects or classes instead.
- Record the pilot's rejection rate. It replaces the guess in line F of the worksheet.
A curve that keeps improving with more sites but not with more hours per site is the clearest signal to buy diversity rather than duration.
How sizing differs by video use case
The right unit changes with the training objective, so match the count to the task before quoting hours.
| Use case | Primary sizing unit | Common sizing mistake |
|---|---|---|
| Action recognition (clip classification) | Labeled clips per class, unique subjects | Buying hours from one camera per class |
| Temporal action segmentation | Fully annotated procedures, step instances per class | Counting unannotated hours as usable |
| Video-language and captioning | Aligned clip-text pairs, verified alignment rate | Trusting ASR narration as captions |
| Video instruction tuning | Question-answer pairs per clip, clip diversity | Many questions on few clips |
| Robot learning from demonstration | Demonstrations per task and object variant | Long sessions with few task resets |
| Evaluation set | Instances per class from unseen sites | Drawing eval from training sites |
For segmentation label specifications, see temporal action segmentation labels. For a domain example where rare errors drive volume, see manufacturing assembly video datasets. Image teams face a parallel problem covered in how many images a vision model needs, and the cross-modality budgeting logic is in how much training data to buy.
Writing the volume into a request
Put counts, not just hours, in the request so suppliers can say whether they can meet them. State minimum unique sites, subjects and cameras; per-class instance targets; the clip specification; what counts as usable; and how the held-out split is drawn. The data request guide gives a template.
Operational footage from real businesses is where the long tail of rare events lives. SourceX sources operational datasets, including new recordings of hands-on work, from US companies on request; it does not hold video in stock, and a request does not guarantee a match. You can describe the scenes, classes and volume you need without naming specific businesses.
Request video data sized to your model
Describe the actions, scenes, per-class targets and usable-hour assumptions your model needs, not the companies you think hold the footage. SourceX looks for US businesses that hold the described data, reviews ownership and consents, and delivers only what the supplying company approves under a license that defines records, uses, term and delivery. Start a buyer request.
Frequently asked questions
What is a reasonable minimum video dataset size for fine-tuning?
There is no universal minimum. For fine-tuning a pretrained video model on a narrow task, diversity of scenes and subjects and a few hundred instances of each class usually matter more than total hours; Kinetics, a broad benchmark built for training from scratch, collected at least 400 clips per class [1]. Confirm with a pilot learning curve.
Should I buy more hours from the same sites or footage from new sites?
Usually new sites. Additional hours from a camera you already have mostly repeat its background and lighting. Buy more of the same only when your learning curve shows a specific rare class is still improving and that class occurs only at those sites.
How do I compare suppliers that quote different units?
Normalize every quote to usable hours, unique scenes, unique subjects and per-class instances. Ask each supplier for a sample and run it through your own rejection pipeline to estimate the usable ratio rather than accepting a stated one.
Sources
- arXiv (Kay, Carreira, Zisserman et al., DeepMind), "The Kinetics Human Action Video Dataset" (2017). https://arxiv.org/pdf/1705.06950
- arXiv (Grauman et al., Ego4D consortium), "Ego4D: Around the World in 3,000 Hours of Egocentric Video" (2021). https://arxiv.org/abs/2110.07058v1
- arXiv (Hestness et al., Baidu Research), "Deep Learning Scaling is Predictable, Empirically" (2017). https://arxiv.org/abs/1712.00409v1
- arXiv, "Smart City Intersections: Intelligence Nodes for Future Metropolises" (2022). https://arxiv.org/pdf/2205.01686
- arXiv (Miech et al.), "HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips" (2019). https://arxiv.org/pdf/1906.03327
- CVPRW 2023 (Hukkelås & Lindseth), "Does Image Anonymization Impact Computer Vision Training?" (2023). https://openaccess.thecvf.com/content/CVPR2023W/WAD/papers/Hukkelas_Does_Image_Anonymization_Impact_Computer_Vision_Training_CVPRW_2023_paper.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.