Skip to content

Video data

Video Question-Answer and Instruction Data for Fine-Tuning Video LLMs

Quick answer

A video instruction tuning dataset pairs clips with questions, instructions and answers so a video LLM learns to reason over time, not just describe frames. Many public video instruction sets are synthetic and generated from captions, so they inherit caption errors and rarely cover operational domains. For SFT that improves temporal and domain reasoning, specify a question taxonomy, require timestamp-grounded answers, verify a human-checked sample, split by source video, and confirm the license covers model-generated content.

By SourceX Editorial · Updated

What separates video QA data from caption data

Video QA and instruction data is a post-training asset with its own taxonomy and grounding rules, while captions are a description asset. A caption says what happens in a clip; a QA pair tests whether the model can answer a specific question about order, cause, count or state at a specific time. Many open pipelines derive QA from captions: a vision-language model writes detailed descriptions from sampled frames, and an LLM then turns those descriptions into open-ended and multiple-choice QA. If you are still sourcing the descriptive layer, start with video-text pairs and dense captions and the dense caption guidelines.

The practical consequence is error propagation. A caption that misses a skipped step, mislabels a tool or compresses two actions into one yields QA pairs that are confidently wrong, and an LLM rewriting step adds fluent phrasing that hides the error. Treat caption-derived QA as a hypothesis that needs human verification on a sample before it enters your SFT mix.

Why public video instruction sets plateau

Public video SFT sets plateau because they over-represent descriptive questions on web video and under-represent temporal, causal and domain reasoning. Academic video QA benchmarks separate descriptive, temporal and causal questions for this reason, and teams commonly see models answer "what is in the scene" far better than "why did this happen" or "what came before." Long videos add a second gap: a model that samples a handful of frames from a 20-minute procedure cannot answer counting or order questions it never saw, and speech or on-screen text often carries the answer.

Benchmark gains from instruction data also deserve scrutiny: always compare the tuned model against the initialized checkpoint on the same frame-sampling settings, or you will credit the data for gains it did not cause. Text SFT research shows curation can beat volume: LIMA reached competitive quality with 1,000 carefully curated examples [1]. For a video LLM, a few thousand verified, grounded, domain-specific pairs can matter more than another hundred thousand synthetic ones.

Question taxonomy to specify before sourcing

Specify the question types you need up front, because suppliers and annotators default to "what is happening" questions. A balanced taxonomy for operational and instructional video usually includes the following, each with a target share of the set:

  • Temporal order: "Which happened first, X or Y?" and "What happened immediately after the valve was closed?"
  • Counting: "How many cartons were placed on the pallet between 00:40 and 01:30?"
  • Causality: "Why did the operator stop the line?" with an answer grounded in a visible event.
  • Next-step prediction: "What is the next step in this procedure?" from a clip that ends mid-task.
  • State change: "Is the panel open or closed at the end of the clip?"
  • Anomaly and safety: "Was lockout applied before the guard was removed?"
  • Procedure compliance: "Which step of the documented procedure was skipped?" which requires a reference SOP; see video-to-SOP step alignment.
  • Abstention: questions the video cannot answer, with a labeled "not visible" response, which helps reduce hallucination; see fine-tuning data that reduces hallucinations.

Domain-expert QA is where synthetic generation falls short. A captioning model cannot reliably say which step a technician skipped or whether a torque sequence was followed; that requires someone who knows the procedure, which is why expert annotations and labels are often the expensive part of a video SFT budget. Source footage for this kind of QA tends to come from procedural and instructional video, manufacturing assembly video and software screencasts.

Record schema with timestamp grounding

Every answer should carry a timestamp or frame range so a reviewer can check it in seconds and you can train or evaluate grounding directly. Without grounding, QA verification means rewatching whole clips, and wrong answers survive review.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "qa_id": "qa_000412",
  "source_video_id": "vid_0187",
  "clip": {"start_s": 312.0, "end_s": 401.5, "fps": 30},
  "question_type": "procedure_compliance",
  "format": "open_ended",
  "question": "Which step of the filter replacement procedure was skipped?",
  "answer": "The operator did not depressurize the housing before removing the cap.",
  "evidence_spans": [{"start_s": 338.2, "end_s": 344.0}],
  "reference_doc": "sop_filter_v3#step4",
  "audio_used": false,
  "generation": {"method": "human_written", "generator_model": null},
  "verification": {"status": "verified", "reviewer_role": "domain_expert", "pass": 1},
  "deid": {"faces": "blurred", "on_screen_text": "redacted", "audio": "removed"},
  "split": "train"
}

Keep generation.method and generator_model on every record. They let you filter synthetic pairs, audit model-output terms and report the human-written share to stakeholders.

Verification, splits and leakage

Verify a random, stratified sample per question type, and split by source video rather than by clip or question. A practical bar is a double-blind check where a second reviewer answers the question from the evidence span, with disagreements adjudicated and the error rate reported per type. Reject batches whose error rate on causal or compliance questions exceeds your threshold, since those are the types most likely to be wrong when derived from captions.

Leakage is the common failure. If clips cut from the same recording land in both train and test, the model memorizes scene, operator and lighting, and your held-out score overstates generalization. Assign split at the source_video_id level, and also hold out whole sites or customers when the deployment target is a new site. Check that none of your evaluation videos appear in public benchmarks you also report on.

Buyer checklist: video SFT and QA data

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckWhat to ask forRed flag
TaxonomyShare of each question type, with examplesOver 70% descriptive "what is happening" questions
GroundingEvidence spans on every answerAnswers with no timestamps
GenerationMethod and generator model per recordSynthetic pairs mixed in without labels
VerificationSampling plan and per-type error rate"Quality checked" with no numbers
SplitsAssignment by source video and siteRandom split at clip or question level
ModalitiesWhether audio and on-screen text are retainedAudio stripped when questions depend on speech
RightsFootage, people, audio and generator termsNo record of consent or model-output terms
PrivacyFace, voice and screen-text handling methodHealth or worker video without a documented method

Rights and provenance for video QA pairs

A video QA dataset carries at least two rights layers: the footage and the QA text, and the QA text may itself be model output. If an LLM or video model generated the questions or answers, check that model's terms of use for restrictions on training competing models, and record the model and version. An audit of 1,800+ public text fine-tuning datasets found license omission above 70% and license error rates above 50% on popular hosting sites [2], so a declared license on a hub page is not proof. The provenance cluster's licensing pages and rights layers in a video clip cover the footage side.

People in operational video raise consent and privacy questions; see recording employees on video for AI datasets. Clinical footage needs HIPAA de-identification by Safe Harbor or Expert Determination, and Safe Harbor's identifier list includes full-face photographs and comparable images [3]; surgical video datasets covers that path.

Where SourceX fits in a video QA data request

SourceX sources operational datasets from US companies on request, including documents, engineering records and new recordings of hands-on work that can underpin domain QA, and manages licensing and ongoing purchases; it does not hold stock, and a request does not guarantee a match. Buyers describe the footage and QA they need, not the companies, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery, with personal details removed or replaced and the method recorded. SourceX does not source scraped web content or generic CCTV, and it does not train models. You can send a video QA data request to SourceX with your taxonomy and record schema. For the wider category, see the video data hub, the AI data hub and domain-specific fine-tuning data.

Request grounded video QA data for your video LLM

Describe the footage, question taxonomy and grounding you need, and SourceX will look for US businesses that hold matching data and assess its licensing permissions. Nothing is contracted until a supplier agrees, and terms are set per deal. Start a buyer request at SourceX.

Sources

  1. Zhou et al., arXiv (NeurIPS 2023), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  2. Longpre et al., Nature Machine Intelligence 6, "A large-scale audit of dataset licensing and attribution in AI" (2024). https://www.nature.com/articles/s42256-024-00888-z
  3. U.S. HHS Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data