Skip to content

Video data

Video-Text Pairs and Dense Video Captions for Video-Language Models

Quick answer

A usable video caption dataset pairs each clip with text that describes what is visible, at the granularity your model needs. That might be one caption per clip, event captions with start and end timestamps, or narration aligned to the steps on screen. You also need written rights to both the footage and the text. Web-scraped pairs fail on both counts: the text is often off-screen chatter, and the licenses are often missing. Buyers who specify the text source, the timestamp format and the authorship chain up front get data they can train on and defend later.

By SourceX Editorial · Updated

Which text sources produce usable video-text pairs?

The text source sets your quality ceiling more than the footage does, so choose it first. Each common source fails in a different, predictable way, and most production corpora mix two or three.

  • Titles, descriptions and alt text. Cheap and plentiful, but written for search and clicks, not for describing frames. Expect brand names, calls to action and no temporal detail.
  • ASR narration transcripts. Abundant in instructional video, but only weakly paired. Speakers ask viewers to subscribe, digress, or describe a step before or after it happens, and ASR adds its own transcription errors. Treat narration as noisy supervision and measure how often the caption describes something actually visible in the clip.
  • Human dense captions. Annotators mark events and write time-bounded descriptions, the format popularized by the ActivityNet Captions benchmark [7]. This text is accurate and temporally grounded, and it costs the most per minute of video.
  • Model recaptions. A captioning model rewrites or densifies text at scale. Check the generating model's output terms before you train on that text, and keep a flag on every recaptioned row so you can audit hallucinations later.
  • Work-record text. Operational video often comes with text a business already writes: step logs, inspection notes, SOP references, ticket comments. That text is domain-accurate but needs time alignment. See our guide to pairing procedure video with written SOPs.

How fine-grained should the captions be?

Match caption granularity to the objective: contrastive pre-training tolerates clip-level text, while dense captioning, temporal grounding and video-LLM tuning need event-level captions with timestamps. Specify granularity in the request, because converting coarse captions to dense ones later means re-annotating.

GranularityTypical unitBest forWatch for
Clip-levelOne caption per 5 to 30 s clipContrastive pre-training, retrievalCaptions that summarize and skip actions
Shot-levelOne caption per camera shotEditing-aware models, scene searchShot-boundary detector errors splitting events
Event-level denseStart, end, caption per event; events may overlapDense captioning, temporal groundingTimestamp drift, missing short events
Step-levelOrdered steps tied to a procedureInstructional and procedural modelsSteps described before or after they occur

Overlapping events matter. Real events run from seconds to minutes and often overlap, so your schema should not force a single non-overlapping segmentation unless the task demands it. For the annotation rules behind each level, see dense video captioning guidelines. For action boundaries without free text, see temporal action segmentation labels.

What does a delivery-ready record look like?

A delivery-ready record names the media file, frame rate and time base, then lists every caption with its timing, author type and rights reference. The fields below are the ones buyers most often regret leaving out.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "clip_id": "vid_000481_c03",
  "media": {"uri": "vid_000481.mp4", "codec": "h264", "fps": 29.97, "duration_s": 94.2, "audio": true},
  "time_base": "seconds_from_file_start",
  "captions": [
    {"start": 3.4, "end": 11.0, "text": "A technician loosens four bolts on the pump housing with a ratchet.",
     "granularity": "event", "author_type": "human_annotator", "language": "en", "lang_variant": "en-US",
     "guideline_version": "dc-v2.1", "review": "second_pass_passed"},
    {"start": 9.8, "end": 22.5, "text": "She lifts the housing cover and sets it on the bench.",
     "granularity": "event", "author_type": "model_recaption", "generator": "recorded_model_id",
     "review": "human_verified"}
  ],
  "asr": {"engine": "recorded_engine_and_version", "file": "vid_000481.vtt", "aligned_to_visual": false},
  "rights": {"footage_license_ref": "L-114", "caption_text_license_ref": "L-114-A",
             "people_consent_ref": "C-22", "music_present": false},
  "deid": {"faces": "not_blurred", "speech_pii": "redacted", "method_log": "deid_run_17"}
}

Three details carry most of the value. Keep time_base explicit, because WebVTT and SRT cues, container edit lists and variable frame rate all shift timestamps. Store ASR separately from visual captions so narration never silently becomes ground truth. Record author_type per caption, because human, ASR and model-generated text need different evaluation and different rights checks. Packaging the dataset card in a machine-readable format such as Croissant, a schema.org-based JSON-LD vocabulary for files and record structure, makes these fields loadable instead of buried in a PDF [4].

Why do rights cover two things in every pair?

Every video-text pair carries at least two rights chains, one for the footage and one for the caption text, and they frequently belong to different parties. A company may own its recorded training sessions while a contractor wrote the captions, or a vendor's model generated them under its own output terms.

In the US, a transfer of copyright ownership is not valid without a signed writing [2]. So "the annotation vendor gave us the captions" is not the same as holding rights to them; for independent contractors, ensure an express assignment of copyright ownership under 17 U.S.C. 204 [2] is executed alongside the agreement. Ask for the annotation agreement with a signed assignment or work-for-hire language that covers the caption text, alongside the footage license. Footage itself can stack further layers (people on camera, music, on-screen displays and logos), covered in rights layers in a video clip.

Public collections do not remove this work. An audit of popular dataset hosting sites found license omission rates above 70% and error rates above 50% [1], so trace any public video-text set to the original terms of both the video host and the caption source. Transparency rules raise the stakes. As of October 2026, general-purpose model providers in the EU must publish a summary of training content under Article 53(1)(d), using the template the Commission issued on 24 July 2025 [5], and California AB 2013 has required generative AI developers serving Californians to post training-data documentation since 1 January 2026 [6]. Both are easier to satisfy when each caption row already says where it came from.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Buyer checklist for a video caption dataset request

Use this checklist when you write a request or evaluate a sample; every item should have a written answer before contract review.

  • Objective: pre-training, retrieval, dense captioning, grounding or video-LLM SFT. For question-answer formats, see video instruction-tuning data.
  • Granularity and timing: clip, shot, event or step; time base; tolerance for boundary error; overlap allowed or not.
  • Text source mix: share of human, ASR and model-generated captions, with the generating model and its output terms named.
  • Language: caption languages and variants, whether translations are human or machine, and domain glossaries (part names, clinical or legal terms).
  • Quality evidence: guideline version, double-annotation rate, agreement measure on a sample, and the hallucination check used for recaptions.
  • Duplicates and leakage: near-duplicate clips and repeated boilerplate captions removed. Text-model research shows duplicated training text raises verbatim copying [3], and repeated intros and outros do the same to captions.
  • Rights: footage license, caption-text authorship and assignment, consents for identifiable people, and music and third-party content.
  • Privacy: faces, voices, names spoken aloud and screens in frame, with the de-identification method recorded.

For licensed footage without captions, see video recordings. For human labeling as a separate input, see expert annotations and labels.

Where captioned operational video comes from

Company video recorded for real work gives you domain footage paired with text the business already wrote, which scraped clips rarely offer. Examples include recorded training sessions, field-service and assembly recordings, and software walkthroughs. Start from the video data hub, then look at corporate training video libraries and software screencasts. If your model generates video instead of describing it, see text-to-video training data. The broader concept is defined under multimodal data.

SourceX sources operational datasets from US companies on request, including new recordings of hands-on work. It does not source scraped web content or generic CCTV, and a request does not guarantee a match. You can describe the video-text data you need on the buyers page using the record fields above.

Request licensed video-text pairs and dense captions

Describe the footage, caption granularity and text source you need. SourceX looks for US businesses that hold that data, reviews ownership and consents, removes or replaces personal details with a recorded method, and delivers under a license that defines records, uses, term and delivery once the supplying company approves the release. Start a buyer request.

Sources

  1. arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  2. Legal Information Institute, Cornell Law School, "17 U.S. Code § 204 - Execution of transfers of copyright ownership". https://law.cornell.edu/uscode/text/17/204
  3. arXiv (Lee et al.), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  4. arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  5. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  6. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  7. ICCV (Krishna et al.), "Dense-Captioning Events in Videos" (2017). https://arxiv.org/abs/1705.00754

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data