Skip to content

Video data

Curating Raw Video into Training Clips: Shot Detection, Filtering and Quality Scoring

Quick answer

A video data curation pipeline turns long, licensed source files into short, single-shot training clips through five stages: shot boundary detection, static and low-motion filtering, text and watermark overlay detection, technical and aesthetic scoring, and captioning. Each clip keeps its source file ID, timestamps and rights record. Ask suppliers for original-quality files, rights metadata and agreed anonymization. Keep segmentation thresholds, scoring and captioning in house, because those choices depend on your model and change between training runs.

By SourceX Editorial · Updated

Why curation decides video model quality

Curation is a training decision, not a cleanup chore, because the same raw pool filtered differently produces measurably different models. The Stable Video Diffusion authors split training into image pretraining, video pretraining and high-quality video fine-tuning, and they treat systematic data curation as a core contribution rather than a preprocessing footnote [1]. Their pipeline removed clips with cuts, little motion, heavy on-screen text and low aesthetic value before pretraining [1].

Panda-70M shows the same pattern for video-language work. Its authors cut about 3.8M high-resolution source videos into roughly 70M semantically consistent clips, captioned each clip with several cross-modality teacher models, and its training split comes to roughly 36 TB once downloaded [2]. At that scale, every filter you can run cheaply on metadata saves expensive decode and GPU time later.

For a buyer, the practical lesson is that raw footage is the input, and the clip index is the product. If you are deciding how much processing to request from a supplier, the owner page on whether AI labs want raw or cleaned data frames the trade-off; this page covers the video-specific stages.

Stage 1: shot boundary detection and clip segmentation

Shot boundary detection should run first, because every later filter assumes a clip shows one continuous camera take. A clip that spans a hard cut teaches a generation model that scenes can teleport, and it confuses temporal action labels.

Two open tools are common. PySceneDetect thresholds frame-to-frame content changes, runs on CPU and is fast, while TransNetV2 is a neural detector that tends to handle gradual transitions such as dissolves and fades better [7]. A common compromise runs the fast detector as a first pass and the neural one as a second pass on ambiguous regions. Stable Video Diffusion went further and ran cut detection at several frame rates to catch cuts that a single pass missed [1].

Segmentation is more than cutting at every detected boundary. Panda-70M splits on shot boundaries and then stitches adjacent pieces back together when they are semantically similar, balancing semantic coherence against clip duration [2]. Typical rules you will set yourself:

  • Minimum and maximum clip duration for your model's context, with long shots split into overlapping windows.
  • A trim of a few frames at each boundary, so encoder artifacts and transition frames do not leak in.
  • Rejection of clips whose boundary confidence falls in an ambiguous band, rather than forcing a guess.
  • Frame-accurate timestamps stored as presentation timestamps (PTS) or timecode, never as "clip 7 of 40".

Failure modes to test on a sample: flash photography and strobe lighting read as cuts, slow pans across uniform scenes hide real cuts, and variable frame rate phone footage shifts timestamps unless you re-time against PTS.

Stage 2: static, low-motion and degenerate clip filters

Motion filtering removes clips that look like video but behave like still images, which otherwise teach a generator to produce frozen output. Stable Video Diffusion computed optical flow scores and dropped clips with little or no motion [1].

The usual implementation computes dense optical flow (Farneback on CPU, or a learned estimator such as RAFT on GPU) on a downsampled, low-frame-rate copy, then takes the mean flow magnitude per clip. Thresholds depend on the domain. A warehouse picking clip with a fixed camera has low global motion but meaningful local motion, so a global mean can wrongly reject it; per-region or percentile statistics work better for fixed-camera footage.

Also filter these degenerate cases, most of which are cheap to detect from decoded frames or container metadata:

  • Black, white or single-color frames, including fade-to-black tails.
  • Slideshow clips where frames repeat for long runs (identical frame hashes).
  • Letterboxing and pillarboxing that waste resolution; crop or reject.
  • Clips with no audio stream when your model expects audio, or audio that is silence.
  • Excessive camera shake or rolling-shutter wobble if your target is stable footage.

Stage 3: text, watermark and overlay detection

Overlay detection removes burned-in text, logos and watermarks, which a generative model will otherwise learn to reproduce and which can create trademark and rights problems. Stable Video Diffusion used OCR-based text detection to drop clips with large amounts of written text [1].

Run a text detector on sampled frames and compute the fraction of frame area covered by text boxes, then threshold per use case. Subtitles and lower thirds sit in predictable bands, so a band-aware rule can crop them rather than reject the whole clip. Static logos and broadcast bugs are better caught by looking for pixels that stay constant across frames while the rest of the image changes, or by a small classifier trained on your own watermark examples.

Overlay findings also feed the rights review. A third-party logo, a stock-footage watermark or a broadcast bug in supposedly first-party footage suggests the supplier may not own every layer of the clip; the guide to rights layers in a video clip walks through footage, people, music and brand layers. Software tutorials are an exception: on-screen text is the signal, so do not apply a generic text filter to screencast datasets.

Stage 4: technical and aesthetic quality scoring

Quality scoring ranks surviving clips so you can choose a pretraining pool and a smaller, higher-quality fine-tuning pool. Stable Video Diffusion scored clips with CLIP-based aesthetic and text-image similarity signals and showed that the curated subset trained better models than an uncurated pool of similar size [1].

Separate technical metrics from aesthetic ones, because they fail differently:

  • Technical: resolution, bitrate per pixel, blockiness and banding from re-encoding, blur or focus score, exposure clipping, interlacing combing, and frame drops. Most are computed from ffprobe metadata plus a few decoded frames.
  • Aesthetic or semantic: learned aesthetic predictors, image-text similarity against a draft caption, and content classifiers for unwanted categories.

Store raw scores, not just pass or fail flags. Thresholds that suit a text-to-video base model will be wrong for an action-recognition model, and you will rerun selection many times. Treat learned aesthetic scores with care: they encode the taste of whoever labeled their training data and can quietly remove industrial, low-light or handheld footage that your model actually needs. ISO/IEC 5259-4 is a useful reference for documenting this as a repeatable data quality process rather than a set of one-off scripts [5].

Stage 5: deduplication and captioning

Deduplication should run after segmentation and before captioning, because duplicate clips waste captioning compute and raise memorization risk. Work on text corpora found that duplicated training data increases verbatim memorization and that deduplication improves models [3]; the same logic applies to reposted, re-encoded or re-cut video. Use clip-level perceptual hashes or embeddings, and see the training data quality hub for near-duplicate detection methods.

Captioning is the last stage and the most model-specific. Panda-70M generated candidate captions from several image and video teacher models and used a retrieval model, fine-tuned on a human-selected subset, to pick the final caption for each clip [2]. Caption formats, density and length are covered in video-text pairs and dense captions and, for generation models specifically, text-to-video training data.

Lineage: tracing every clip back to its source and license

Every clip record should carry enough lineage to find its source file, its exact time range and the license that covers it. Without this, you cannot honor a takedown, remove one supplier's footage, or answer a diligence question about where a generated output's training data came from. Audits of public datasets found license information frequently missing or wrong once data had been copied and re-hosted [4], and the Panda-70M release itself tells users to follow the related license for the video samples [2].

Lineage also prevents evaluation leakage. Split train, validation and test sets by source video (and ideally by supplier, location or camera), not by clip. Adjacent clips from one recording share lighting, people and background, so a clip-level random split puts near-copies on both sides and inflates test scores.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "clip_id": "clp_000418_s03_w02",
  "source_file_id": "src_000418",
  "source_sha256": "9f2c...e71a",
  "supplier_batch": "batch_2026_07",
  "license_ref": "lic_ref_0042",
  "pts_start": 128.512,
  "pts_end": 134.017,
  "shot_detector": "transnetv2+pyscenedetect",
  "boundary_confidence": 0.94,
  "flow_mean": 3.71,
  "text_area_frac": 0.02,
  "watermark_flag": false,
  "aesthetic_score": 5.4,
  "blur_score": 0.12,
  "anonymization": "faces_blurred_v2",
  "dedup_cluster": "dc_88213",
  "split": "train",
  "caption_model": "captioner_v3"
}

Package this index as a sidecar table alongside the media; video dataset packaging and sidecar metadata covers containers and clip indexes, and a dataset card with license, size and provenance fields is a familiar way to publish it internally [6].

What to ask suppliers to do versus keeping in house

Suppliers should deliver original-quality footage with rights and anonymization handled; your team should own anything that encodes a modeling choice. Supplier-side trimming that discards "boring" footage removes exactly the static and edge-case material you may want for negative examples or robustness.

Illustrative example: invented to show structure; it does not describe an available dataset.

StageSupplier sideBuyer sideWhy
Original files and codecsDeliver originals, no re-encodeVerify hashes and ffprobe outputRe-encoding destroys quality signals
Rights and consent metadataProvide per-file rights recordMap to license reference per clipOnly the holder knows who appears and what was agreed
Anonymization (faces, plates, screens)Apply agreed method and record itSpot-check a samplePersonal data should not leave the supplier unprotected
Removal of out-of-scope filesRemove files the license excludesConfirm against manifestScope belongs in the license
Shot detection and segmentationOptional coarse markersFinal boundaries and windowsDuration rules are model-specific
Motion, overlay and quality filtersNoYes, with stored raw scoresThresholds change per run
DeduplicationRemove exact file duplicatesClip-level near-duplicatesNeeds your full corpus to compare
CaptioningSupply existing human descriptions if anyModel captionsCaption style is a model choice

If footage shows faces, voices or workers, align the anonymization method before delivery; the privacy hub and the guide to biometric data rules such as BIPA cover what reviewers will ask.

Where SourceX fits in a video curation workflow

SourceX sources operational data, including new recordings of hands-on work, from US companies on request; it does not hold stock, and a request does not guarantee a match. It does not source scraped web content or generic CCTV. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Describe the footage, file formats and metadata your pipeline expects on the buyer request page, or browse the owner page on licensing video recordings and the video data hub.

Request licensed raw video for your curation pipeline

SourceX looks for US businesses that hold the footage you describe, with every release approved by the supplying company and allowed uses set in a license. Nothing is contracted until a supplier agrees, and SourceX does not train models. Describe the video data your pipeline needs at sourcex.si/buyers.

Sources

  1. arXiv (Blattmann et al., Stability AI), "Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets" (2023). https://arxiv.org/abs/2311.15127
  2. Snap Research (GitHub), "Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers (official repository)" (2024). https://github.com/snap-research/Panda-70M
  3. arXiv (Lee et al.), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  4. arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  5. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Data quality process framework" (2024). https://www.iso.org/standard/81093.html
  6. Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
  7. arXiv (NVIDIA), "Cosmos World Foundation Model Platform for Physical AI". https://arxiv.org/pdf/2501.03575

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data