Video data
Training Data for Text-to-Video Models: What Video Generation Teams Look For
Quick answer
Text-to-video teams need far more than "lots of video." Good training data is footage cut into clean single-shot clips. It passes filters for resolution, motion, aesthetics, burned-in text and watermarks, and each clip carries captions that name subject, action, camera movement, lighting and style. It also needs clear rights covering the footage, the people on screen and any do-not-train signals. Public collections rarely meet the rights bar, so licensed or audited sources matter more for generation than for recognition tasks.
By SourceX Editorial · Updated
Why generation data is judged differently from understanding data
Generation models learn to reproduce pixels, so every defect in the training set can surface in outputs. A classifier trained on shaky phone video with a broadcaster logo still learns "forklift." A diffusion model trained on that clip learns to paint the logo and the shake. That is why the general advice in the video data hub and the dataset quality guide needs a stricter, generation-specific layer.
Published video diffusion recipes commonly run in stages: image pretraining, video pretraining on a large filtered corpus, then fine-tuning on a much smaller high-quality set. Curation decisions made at the pretraining stage carry through to the final model, so they are not just a cost-saving step. For buyers, the practical result is two different shopping lists: broad, filtered volume for pretraining, and narrow, top-quality clips for fine-tuning.
Quality filters that decide whether a clip is usable
A clip is usable for generation when it is one continuous shot with real but controlled motion, adequate resolution and no overlaid text or marks. Most production pipelines run these checks automatically and then sample by hand. Ask any supplier to report the same metrics so you can filter before you pay for transfer.
- Shot boundaries. Run a scene-cut detector such as PySceneDetect, or an embedding-similarity splitter, and keep clips within one shot. Splitting before captioning keeps each caption tied to one coherent shot.
- Motion amount. Compute mean optical-flow magnitude (for example RAFT or Farneback) per clip. Near-zero flow means a static shot that teaches the model to make slideshows. Very high flow often means camera shake or fast pans.
- Aesthetic and technical quality. Use an aesthetic predictor (a LAION-style CLIP head) and a no-reference quality score such as DOVER for blur, noise and compression artifacts.
- Overlaid text and watermarks. OCR-based text-area ratio and a watermark classifier. Burned-in subtitles, lower-thirds, timestamps and channel bugs are among the most common contaminants.
- Technical metadata. ffprobe fields: width, height, avg_frame_rate, codec_name, bit_rate, duration, plus letterbox detection. Interlaced or variable-frame-rate sources need normalizing.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Filter | Pretraining threshold (example) | Fine-tuning threshold (example) | Common failure it removes |
|---|---|---|---|
| Resolution | short side >= 360 px | short side >= 720 px, native not upscaled | Blur, upscaling artifacts |
| Frame rate | >= 15 fps after normalization | constant 24-30 fps source | Judder, duplicated frames |
| Clip length | 2-20 s, single shot | 4-12 s, single shot | Hard cuts inside a sample |
| Mean optical flow | above static floor, below shake ceiling | middle band, smooth camera | Slideshows, shaky handheld |
| Aesthetic score | above corpus median | top decile | Badly lit, poorly framed footage |
| Text-area ratio | < 5% of frame | ~0% | Subtitles, logos, timestamps |
| Watermark classifier | below threshold | zero detections on manual sample | Stock watermarks reproduced in outputs |
Record the thresholds and score versions in your data card. When a model later produces a recurring artifact, you can trace it back to the filter that should have caught it.
Captions that teach camera, motion and style
Captions for generation should describe what a prompt writer would ask for: subject, action, setting, camera motion, framing, lighting and visual style. Short alt-text captions ("a man in a kitchen") leave the model unable to follow "slow dolly-in, low-angle, warm tungsten light." Separate camera vocabulary from subject vocabulary so you can test prompt adherence on each.
Captioning at scale usually means models, not people, and a single captioner is rarely enough. Large public caption sets have combined several captioning models with titles, subtitles and sampled frames, then had humans rate a sample, because each model fails on different content. Supplier-side metadata, such as shot lists, production notes or procedure names, is valuable input to that process. For timestamped, event-level captions and hallucination checks, see dense video caption guidelines and video-text pairs for video-language models.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"clip_id": "clp_000412",
"source_file_sha256": "9f2c...e1",
"start_s": 12.40, "end_s": 19.90,
"width": 1920, "height": 1080, "fps": 29.97,
"filters": {"mean_flow": 3.1, "aesthetic": 6.2, "text_area_ratio": 0.0, "watermark_p": 0.01},
"caption_short": "A technician tightens a valve on a steel pipe.",
"caption_long": "Close-up of gloved hands turning a red valve wheel on a horizontal steel pipe; slow push-in from the right; overhead fluorescent light; documentary style, shallow depth of field.",
"camera": {"motion": "dolly_in", "angle": "eye_level", "shot_size": "close_up"},
"people_visible": "hands_only",
"audio": "removed",
"rights": {"license_id": "LIC-xxxx", "do_not_train_flag": false, "release_status": "workplace_notice_on_file"}
}
Rights checks matter more for generation
Generation models can reproduce what they saw, so the rights in every layer of a clip carry more weight than they do for recognition models. Do not rely on license labels from public collections. The Data Provenance Initiative found license omission above 70% and license error rates above 50% on popular dataset hosting sites [1]. Licensed footage with a documented chain of title, compared in licensed vs synthetic vs scraped data, is the safer base for a model you plan to ship.
A single clip can carry separate rights in the footage, the people shown, background music, brands, artwork and screen content. Each layer is broken down in rights layers in a video clip. Check machine-readable training preferences too. Content credentials can carry a creator's do-not-train preference [2], and the C2PA ecosystem defines training and data mining assertions for this purpose, now specified by the Creator Assertions Working Group [3]. These signals state a preference rather than technically blocking use, so your ingestion pipeline should read these manifests and drop or quarantine flagged files. If you need footage with a documented rights record behind every clip, you can describe it to SourceX as a buyer.
Regulatory expectations point the same way. As of October 2026, providers of general-purpose AI models placed on the EU market must keep a copyright policy that honors Article 4(3) rights reservations [4] and publish a training-content summary using the AI Office template dated 24 July 2025 [5]. In the US, the Copyright Office's Part 3 report on generative AI training is still a pre-publication version [6]. Keep per-source license records that can feed both a data card and a public summary.
Likeness, memorization and output-similarity risk
Real people in training video create risk at output time, not only at collection. Carlini et al. extracted more than a thousand training images from image diffusion models, including photos of individual people and trademarked logos [7]. Video models have not been shown to be immune. A clip of a recognizable face repeated across a small fine-tuning set is a likeness risk if prompts can steer the model to regenerate it.
Practical controls: deduplicate near-identical clips before training, cap repeats per source, and blur or crop bystanders (see video anonymization for AI training). Where people are the subject, ask for releases or documented workplace notice, as covered in recording employees on video for AI datasets. Style-specific fine-tuning carries its own exposure. Commentary on Japan's draft guidance under Article 30-4 notes that training aimed at outputs reproducing a particular creator's expression may fall outside the exception [8]. Run output-similarity tests against held-out source clips before release.
Matching sources to pretraining and fine-tuning
Pretraining needs breadth across subjects, environments and camera styles, while fine-tuning needs a small set of near-perfect clips that match your target look. Operational footage from businesses is rarely broadcast-grade, but it brings real-world motion that stock libraries underrepresent. Examples include hands-on work, equipment, facilities and procedures, recorded with consistent cameras and known consent. Purpose-shot recordings can be specified for fine-tuning: fixed frame rate, 4K source, scripted camera moves, and audio removed or kept per your plan.
Teams planning very large pretraining runs should read sourcing for foundation-model pre-training teams. For packaging clips, captions and filter scores together, columnar formats such as Parquet with Arrow readers keep metadata queryable (Arrow-based dataset formats). The video recordings page describes what kinds of licensed recordings can be requested.
Request checklist for text-to-video data
A good request specifies the clip, the captions and the rights record, not only hours of footage. Use this as a starting template.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | What to specify | Why it matters |
|---|---|---|
| Use | pretraining, fine-tuning, image-to-video, editing | Sets volume versus quality trade-off |
| Subjects and settings | e.g., industrial maintenance, food prep, retail interiors | Coverage gaps show up as prompt failures |
| Technical floor | resolution, fps, codec, min clip length, single shot | Avoids upscaled or variable-frame-rate sources |
| Filter metrics delivered | flow, aesthetic, text-area, watermark scores per clip | Lets you re-threshold without re-downloading |
| Caption spec | short + long caption, camera fields, style fields | Drives prompt adherence |
| People on screen | none, hands only, consenting subjects, bystanders blurred | Likeness and privacy exposure |
| Rights record | footage ownership, releases, music and brand handling, C2PA/do-not-train flags | Supports copyright policy and training summaries |
| Allowed uses | training, commercial deployment of outputs, term | Must match how the model will ship |
Request licensed video for text-to-video training
SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and does not source scraped web content or generic CCTV. Every dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery, and release requires the supplying company's approval and a license that defines records, uses, term and delivery. A request does not guarantee a match; you describe the data, not the businesses. Describe the video you need on the SourceX buyers page.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Sources
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- MIT Technology Review, "Adobe wants to make it easier for artists to blacklist their work from AI scraping" (2024). https://www.technologyreview.com/2024/10/08/1105234/adobe-wants-to-make-it-easier-for-artists-to-blacklist-their-work-from-ai-scraping
- Coalition for Content Provenance and Authenticity (C2PA), "C2PA clarification to C2PA TDM assertions reference". https://c2pa.org/c2pa-clarification-to-c2pa-tdm-assertions-reference/
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- USENIX Security 2023 (Carlini et al.), "Extracting Training Data from Diffusion Models" (2023). https://www.usenix.org/conference/usenixsecurity23/presentation/carlini
- Privacy World, "Japan's new draft guidelines on AI and copyright: is it really OK to train AI using pirated materials?" (2024). https://www.privacyworld.blog/2024/03/japans-new-draft-guidelines-on-ai-and-copyright-is-it-really-ok-to-train-ai-using-pirated-materials/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.