Skip to content

Video data

Long-Video Understanding Data: Hour-Scale Recordings with Temporal Questions

Quick answer

Long-video understanding data is continuous, untrimmed recording, typically 30 minutes to a full shift, paired with questions that can only be answered by remembering events far apart in time: what changed, how many times, in what order. As of October 2026, the most widely used public benchmarks evaluate minutes to about an hour, so teams training long-context video LLMs usually need licensed operational footage plus a timestamped ground-truth source, ideally a system event log, and splits that hold out whole recordings rather than clips.

By SourceX Editorial · Updated

Why short-clip corpora do not train long-range memory

Clip duration is a poor proxy for how much temporal context a question actually needs. EgoSchema made this point with "temporal certificate sets", the minimum footage a human must watch to verify an answer, and reported intrinsic temporal lengths far longer than earlier video QA datasets, yet its clips are three minutes long [1]. Later work notes that EgoSchema's definition of "long" is those three-minute videos, which is why newer benchmarks push to much longer egocentric recordings [2].

LongVideoBench goes further, with 6,678 multiple-choice questions on 3,763 web-collected videos in duration groups up to 15 to 60 minutes, and reports that open-source models do not scale properly as more frames are added [3]. Video-MME similarly reports results by short, medium and long subsets [4]. The pattern for buyers: widely used evaluation sets top out around an hour, with only newer benchmarks pushing to multi-hour recordings [2], while training corpora with matching continuity and question density are scarce, and web-collected sources often carry non-commercial terms. Check the license on any benchmark before it touches a training run.

The scarce input is not more footage but continuity. A 2-hour recording cut into 120 one-minute clips loses exactly the cross-segment dependencies a long-context model must learn. Large egocentric corpora such as Ego4D (3,670 hours from 931 camera wearers) show that hour-scale collection with consent and de-identification is feasible [5], but most of its footage is daily-life activity rather than the work processes enterprise models are asked to follow. For short-clip QA and instruction formats, see video question-answer and instruction data; this page covers what changes when the unit is an hour.

Question types that actually require long-range memory

A long-range question is one whose evidence spans two or more segments separated by minutes, so no single window of frames contains the answer. Write the taxonomy before you source, because it determines which recordings and which ground-truth fields you need.

Question typeExample templateMinimum evidence spanGround truth needed
State change since"What was different about station 3 at the end of the shift versus the start?"Start and end of recordingAsset state snapshots or start/end log entries
Counting over time"How many times was the reject bin emptied?"Whole recordingDiscrete event log with timestamps
Ordering"Was the calibration check done before or after the first batch?"Two distant eventsOrdered event IDs
Re-identification of objects"Is the pallet loaded at 14:20 the one scanned at 13:05?"Two sightingsObject IDs (barcode, work order)
Causal recall"Why did the line stop the second time?"Cause and effect, often far apartIncident or downtime reason codes
Summarization"Summarize the shift in five steps."Whole recordingStep or phase segmentation
Needle retrieval"When did the operator first open the manual?"One short moment in a long spanPrecise timestamp, plus distractors

Distribute evidence positions deliberately. Language models degrade when relevant information sits in the middle of a long context [6], and a QA set whose answers cluster in the first and last five minutes will overstate memory. Record, per question, the timestamps of every evidence segment and the gap between them, so you can report accuracy by evidence span and by position.

Operational event logs as timestamped ground truth

The cheapest reliable labels for hour-scale video are often the systems that already recorded what happened. A manufacturing execution system, warehouse management scan log, point-of-sale journal, ticketing system or machine PLC history logs discrete events with timestamps, object IDs and reason codes, and those fields answer counting, ordering and state-change questions without a human watching the full hour.

Ask for logs in a structured exchange format. OCEL 2.0 defines an object-centric event log metamodel with SQLite, XML and JSON exchange formats [7], which maps well to video because one event can touch several objects (an operator, a work order, a fixture). Whatever the format, require these fields: event ID, event type, timestamp with timezone and source clock, object IDs, actor role (not name), and the camera IDs covering that location.

Failure modes to plan for:

  • Clock drift. Camera NVR clocks and MES servers drift apart by seconds to minutes. Require a sync method (NTP logs, a visible clock in frame, or a known event marked in both) and record the measured offset per recording.
  • Off-camera events. A scan logged at a station outside the field of view yields unanswerable questions. Map each event type to camera coverage before generating QA.
  • Logged but not visible. Software events (a status flag change) may have no visual correlate. Tag events as visual, partially visual or non-visual.
  • Log gaps. Systems go offline. Request the system's uptime record so missing events are not read as "zero occurrences".

Event logs also carry their own privacy risk: sequences and timestamps can make individuals unique even without names [8]. Replace operator IDs with stable pseudonyms consistent between video metadata and the log, and review timestamp precision before delivery. See the privacy and de-identification guides for methods.

A request specification for hour-scale recordings

Specify long-video data by recording continuity, ground-truth availability and question coverage, not by total hours alone. A request that says "500 hours of warehouse video" invites 30,000 one-minute clips.

Illustrative example: invented to show structure; it does not describe an available dataset.

request: long_video_understanding
recording_unit:
  min_continuous_duration_min: 45      # untrimmed, no cuts
  max_gap_allowed_sec: 5               # NVR dropouts tolerated
  target_unit: full_shift              # or task_session, meeting
  views: [fixed_overhead, station_cam] # egocentric optional
  fps_min: 15
  resolution_min: 1280x720
  audio: optional_redacted
ground_truth:
  event_log_required: true
  log_format: [OCEL2_json, csv]
  fields: [event_id, event_type, ts_utc, clock_source, object_ids,
           actor_role, camera_ids, visual_flag]
  clock_offset_measured: true
questions:
  types: [state_change, count, order, object_reid, causal, summary, needle]
  min_evidence_gap_min: 10
  per_hour_target: 20
  answer_format: [multiple_choice_5, free_text_with_timestamps]
splits:
  unit: recording
  group_by: [site, camera_id, week]
privacy:
  faces: blur_or_retain_per_license
  screens_and_documents: review
  log_pseudonymization: stable_per_person
delivery:
  container: mp4_h264_or_h265
  sidecar: per_recording_json
  sharding: webdataset_tar

The worked record below shows the per-recording sidecar a buyer should expect, linking video, log and questions by shared IDs:

Illustrative example: invented to show structure; it does not describe an available dataset.

{"recording_id": "rec_0412", "site_pseudo": "site_B", "camera_id": "cam_07",
 "start_utc": "2026-03-11T13:00:02Z", "duration_s": 14390, "clock_offset_s": -1.8,
 "events_ref": "logs/rec_0412.ocel.json",
 "questions": [{"q_id": "q_0412_09", "type": "count",
   "text": "How many pallets left dock 2 before the first forklift battery swap?",
   "evidence": [[1820, 1844], [6210, 6233], [9905, 9930]],
   "answer": "2", "derived_from": ["ev_118", "ev_342"]}]}

Splits and leakage at the recording level

Hold out entire recordings, and usually entire sites or cameras, never clips from the same recording. Adjacent windows from one shift share lighting, layout, people and ongoing tasks, so a clip-level random split lets the model answer test questions from memorized context.

Near-duplicate content inflates scores even when files differ: deduplication research on text corpora found heavy repetition and showed it increases memorized output [9]. Video has its own versions: the same fixed camera on consecutive days, re-encoded copies, and overlapping fields of view from two cameras covering one aisle. Group splits by site, camera and time block; run perceptual-hash or embedding similarity across splits; and keep question templates out of training if they also appear in your evaluation set.

If you also evaluate on public benchmarks, check whether their source videos could have entered your pretraining via web scrapes. Licensed operational footage from private sites is less likely to overlap with public web video, which is one reason teams seek it for held-out evaluation.

Storage, delivery and review effort scale with hours

Every cost in a long-video project scales with continuous hours, not with clip count. One hour of 1080p H.264 at a typical 5 to 8 Mbps is roughly 2 to 4 gigabytes, and a few thousand hours is a multi-terabyte delivery before audio, sidecars or alternative encodings.

Plan the package early. WebDataset tar shards, often about 1 GB each, are a common way to stream multi-terabyte video corpora into training [10]; keep one recording per shard group so splits stay clean, and index each recording with offsets. See packaging video datasets for container and sidecar conventions, and offline transfer appliances when network transfer is impractical.

Privacy review effort also scales. An hour of footage can include dozens of faces, badges, screens with customer records, whiteboards and overheard conversation; reviewers sample by hour, not by clip. Workplace footage raises employee notice, consent and audio-recording issues covered in recording employees on video for AI datasets, and the rights layers in a video clip guide covers music, brands and on-screen content.

Where hour-scale operational video comes from

Continuous footage with matching logs exists where businesses already record work for safety, quality or training: fixed station cameras in assembly and packing, dock and yard cameras, body-worn or head-mounted cameras in field service, and recorded meetings or screen sessions. Related pages cover manufacturing assembly video, warehouse operations video, developer session recordings and multimodal meeting recordings; the video data hub compares sources. For step-level labels inside long recordings, see temporal action segmentation labels.

SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and manages the licensing process; nothing is held in stock and a request does not guarantee a match. It does not source scraped web content or generic CCTV. If your spec looks like the template above, you can describe the long-video data you need, along with licensed video recordings and egocentric video of skilled manual work.

Diligence questions before you license long recordings

Ask these before signing, because they are hard to fix after delivery:

  1. What is the distribution of continuous recording lengths, and how are dropouts recorded?
  2. Which system logs cover the same period, at what clock precision, and with what measured offset?
  3. Which event types are visible on camera, and which locations are uncovered?
  4. Who appears in the footage, what notice or consent applies, and is audio included?
  5. What is removed or replaced (faces, badges, screens, names in logs), how, and was a sample checked?
  6. Are recordings from the same site or camera kept together for split assignment?
  7. Do the license terms cover training, evaluation and derived question sets, for what term?

For a broader checklist, see questions to ask a training data vendor.

Request hour-scale video for long-context models

SourceX looks for US businesses that hold the continuous recordings and logs you describe, rights-reviews each dataset for ownership and consents, and delivers under a license defining records, uses, term and delivery once the supplying company approves the release. Personal details are removed or replaced before delivery through private, access-controlled workflows. Start a long-video data request.

Sources

  1. NeurIPS 2023 (Mangalam, Akshulakov, Malik), "EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding" (2023). https://neurips.cc/virtual/2023/poster/73630
  2. arXiv, "X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding" (2025). https://arxiv.org/html/2501.06835v2
  3. arXiv (Wu, Li, Chen, Li; NeurIPS 2024), "LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding" (2024). https://arxiv.org/abs/2407.15754v1
  4. arXiv, "Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis" (2024). https://arxiv.org/pdf/2405.21075
  5. arXiv (Grauman et al., Ego4D consortium), "Ego4D: Around the World in 3,000 Hours of Egocentric Video" (2021). https://arxiv.org/abs/2110.07058v1
  6. arXiv (Liu et al.), "Lost in the Middle: How Language Models Use Long Contexts" (2023). https://arxiv.org/pdf/2307.03172
  7. OCEL standard authors (ocel-standard.org), "OCEL (Object-Centric Event Log) 2.0 Specification" (2023). https://arxiv.org/pdf/2403.01975
  8. arXiv (Nunez von Voigt et al.; CAiSE 2020), "Quantifying the Re-identification Risk of Event Logs for Process Mining" (2020). https://arxiv.org/pdf/2003.10707
  9. arXiv (Lee et al.; ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  10. Hugging Face, "WebDataset (Hub documentation)". https://huggingface.co/docs/hub/datasets-webdataset

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data