Video data
Writing Dense Video Captioning Guidelines: Granularity, Timestamps and Hallucination Checks
Quick answer
Dense video captioning guidelines must fix three things before any annotator starts: what counts as an event and what each caption must describe, how start and end timestamps are placed when events overlap, and which inferences are banned because the frames cannot support them. Pair those rules with layered QA (qualifier tasks, gold clips, multi-annotator agreement and event-level hallucination audits) and disclose whenever a model drafted the captions that humans edited.
By SourceX Editorial · Updated
Dense captioning, as defined in the work that introduced ActivityNet Captions, means detecting the events in a video and describing each one in natural language [1]. That definition hides most of the hard decisions. "Event" has no natural unit, "describe" has no natural depth, and video-LLMs trained on loose captions learn to produce fluent text about things that never appeared on screen. This page is a guideline template for teams commissioning or auditing caption annotation. For finding captioned footage itself, see video-text pairs and dense captions; for the wider video landscape, start at the video data hub.
Define the event unit before you define caption length
The event definition controls granularity more than any word-count rule does. Without it, one annotator writes "a technician repairs a pump" for a four-minute clip and another writes forty captions, one per bolt.
Pick one primary unit and write it into the spec:
- Goal-level events (task-sized, often 10 to 120 seconds): useful for summarization and long-video QA.
- Step-level events (one procedural step, such as "removes the gasket"): the default for procedural and instructional footage, and the unit that aligns with step-annotated procedural video and video-to-SOP step alignment.
- Atomic actions (a single verb-object pair, often under 3 seconds): usually better captured as structured labels than as prose; see temporal action segmentation labels.
Then set a minimum event length (for example 1 second, or a frame count at the source frame rate) and a merge rule: two consecutive actions with the same verb and object and a gap under a set threshold become one event. State whether a paragraph-level summary caption for the whole clip is also required, because many video-LLM SFT mixes want both levels.
Specify what every caption must describe, in a fixed order
A caption checklist with a fixed slot order produces more consistent and more auditable text than "describe what you see." Annotators who follow an order skip fewer elements, and reviewers can check each slot.
The slots most teams need are:
- Actor and action: who (by visible role or appearance, not identity) does what, with a concrete verb. "Picks up," not "handles."
- Objects and state changes: the object, its relevant attributes, and its before and after state ("the empty tray is now filled").
- Spatial relations: where the action happens relative to stable landmarks ("on the left conveyor").
- Camera behavior: pan, zoom, cut, handheld shake or a static shot. Generation teams need this; see text-to-video training data.
- On-screen text: transcribe it verbatim in quotes if legible, or mark it
[illegible]. Never paraphrase a label or error code. - Salient audio events: alarms, speech onset or machine sounds, tagged as audio so they can be dropped for vision-only training.
Specify tense (present simple), voice (active), vocabulary (a controlled list of verbs and part names if the domain has one) and a hard ban on meta-language such as "the video shows." Fine-grained object and hand detail needs its own conventions; hand-object interaction annotation covers grasp and contact vocabulary.
Set timestamp rules that survive overlapping events
Timestamp rules must define boundaries by observable change, a tolerance, and an explicit policy for overlap. Boundary disagreement is normal; the guideline's job is to make it small and measurable.
Write these rules down:
- Start is the first frame where the action's motion begins (the hand starts moving toward the object), not when the object is first visible.
- End is the first frame where the resulting state is stable, or the action's motion stops.
- Tolerance: annotators should match a reference within a stated window (for example ±0.5 seconds for step-level events). Calibrate it on pilot data rather than guessing.
- Overlap: events may overlap; each gets its own caption and its own interval. Do not split a long background event to make room for a short foreground one.
- Format: seconds as decimals with the source frame rate recorded, or
HH:MM:SS.mmm, plus the video's duration so offsets can be validated. - Gaps: idle stretches are either left uncaptioned or captioned as "no task activity," and the spec says which.
Measure boundary quality with interval metrics rather than by eye. Action segmentation work reports segmental F1 at 10%, 25% and 50% IoU overlap together with an edit score [2]; the same thresholds give caption audits a common language for "how close is close enough."
Ban inferences the frames cannot support
The prohibited-content list is the main defense against hallucination at the source. Most fabricated detail in human captions comes from annotators filling gaps with plausible guesses.
Ban, with examples in the guideline:
- Identity: no names, employers or demographic labels unless printed on screen and in scope; use roles ("the operator").
- Emotion and intent: no "frustrated," "carefully" or "trying to" unless a visible cue (a gesture, a spoken statement) supports it. Write the cue instead.
- Off-screen causes and outcomes: no "after the machine failed" if the failure is not shown.
- Counts and measurements that cannot be verified from frames ("about 20 boxes" when only part of the pallet is visible).
- Brand or model names read from blurry text.
Offer an escape hatch: [uncertain: ...] tags let annotators flag what they think they see without asserting it. Reviewers then confirm or delete. Two failure forms matter: wrong details about content that is present (a red valve called blue) and content that never appeared at all. Your banned list should target both, and the second is the harder one to catch in review.
Practical artifact: guideline excerpt and caption record
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"video_id": "vid_000417",
"duration_s": 184.20,
"fps": 30,
"guideline_version": "dvc-spec-2.3",
"event_unit": "step",
"summary_caption": "A technician replaces a filter cartridge on a wall-mounted unit and restarts it.",
"events": [
{
"event_id": "e01",
"start_s": 12.40,
"end_s": 21.07,
"caption": "The technician turns the gray valve handle a quarter turn clockwise; the pressure gauge needle drops to zero.",
"slots": {"camera": "static", "on_screen_text": ["\"CLOSE BEFORE SERVICE\""], "audio": []},
"uncertain": []
},
{
"event_id": "e02",
"start_s": 19.80,
"end_s": 47.53,
"caption": "The technician unscrews the clear housing and lifts out a brown-stained cartridge.",
"slots": {"camera": "slow zoom in", "on_screen_text": [], "audio": ["short hiss at 20.1 s"]},
"uncertain": ["[uncertain: label on cartridge, illegible]"]
}
],
"provenance": {"draft_source": "model", "draft_model": "internal-captioner-v4", "human_edit": true,
"edit_distance_ratio": 0.38, "annotator_id": "ann_22", "reviewer_id": "rev_05"}
}
Note the overlap between e01 and e02, the verbatim quoted text, the uncertainty tag instead of a guessed label, and a provenance block that records whether a model drafted the caption.
Build QA in layers, each catching a different error
No single check catches hallucinated captions, so QA should stack screening, embedded gold, agreement and an event-level audit. The layered pattern below adapts practice from large video segmentation projects, which relied on annotator screening, quality checks and repeated review rounds [3].
QA layers checklist
| Layer | What it catches | How to run it |
|---|---|---|
| Qualifier task | Annotators who misread the spec | 5 to 10 reference clips; pass on slot coverage, banned-inference count and boundary tolerance |
| Gold clips | Drift over time | Mix reference clips into live queues; track per-annotator scores weekly |
| Overlap sample | Inconsistent granularity | Double-annotate a share of clips; compare event counts and boundary IoU |
| Event-level audit | Missing, wrong and invented events | Reviewer marks each reference event as covered, incorrect or missing, and each caption claim as supported or not |
| Text checks | Format and vocabulary errors | Automated lint for banned words ("seems," "trying"), meta-language, timestamp order and bounds |
Agreement statistics need care with free text. Choose a metric that fits the label type, such as Cohen's kappa for two raters on categorical judgments or Krippendorff's alpha for many raters and missing data [4]; apply it to the structured judgments (event present, claim supported, boundary within tolerance), not to raw caption strings. Expect residual errors even after QA: audits of widely used test sets found an average label error rate of at least 3.3% [6], and a caption benchmark with similar noise can reorder models.
Score hallucination at the event level, not with n-gram metrics
Event-level precision and recall measure what buyers care about: whether the caption covers what happened and invents nothing. N-gram metrics such as CIDEr reward word overlap and are poorly suited to long, detailed descriptions.
Recent detailed-captioning evaluations extract events from each caption, match them against reference events, and count missing, incorrect and invented ones. For your own deliveries, run the same logic on an audit sample:
- Event recall = reference events covered ÷ reference events.
- Claim precision = supported claims ÷ total claims in captions.
- Hallucination rate = unsupported or contradicted claims ÷ total claims, reported separately for on-screen text, counts and attributes.
If an LLM judge does the extraction, spot-check it with humans, since judge errors are a second source of noise.
Disclose model-drafted captions and measure the edit rate
When models draft captions and humans edit them, the dataset is partly synthetic, and buyers should know which parts. Large public corpora already rely on model teachers; Panda-70M, for example, captioned its clips with multiple cross-modality teacher models [5].
Require per-caption provenance (draft_source, model name or version, human_edit flag) and a normalized edit distance between draft and final text. A very low edit rate can mean the model was good, or that reviewers rubber-stamped it; check it against hallucination rates on the audit sample. Note that model drafting also imports that model's vocabulary and blind spots, which matters if you train a competitor model or use the captions for evaluation. Document all of this in a dataset card covering collection, annotation method and intended use [7], and link guideline versions to delivered files using the conventions in delivery formats and schemas.
Questions to send a caption vendor or data supplier
Ask these before you sign, and keep the answers with the dataset card:
- What is the event unit, minimum event length and overlap policy, and can we see the guideline version used?
- Which slots are mandatory, and how are on-screen text and audio handled?
- What share of captions was model-drafted, and what was the median edit rate?
- What were the qualifier pass threshold, gold-clip scores and boundary agreement at the stated tolerance?
- What were event recall and hallucination rate on an independent audit sample?
- Who appears in the footage, and what consent and rights cover the video, audio and any on-screen content? See rights layers in a video clip and recording workers on video.
For general annotation terms, the data annotation glossary entry and the page on licensing expert annotations and labels cover the commercial side. If guidelines are shared across modalities, compare with image annotation guidelines and label taxonomies and SFT demonstration writing guidelines.
Where captioned operational video comes from
Real operational video, such as workers performing procedures, is held by businesses rather than public sites, and SourceX sources it on request from US companies, including new recordings of hands-on work. Nothing is held in stock and a request does not guarantee a match; you describe the footage and the captioning you need on the buyer request page, and every release is approved by the supplying company.
Request captioned operational video
SourceX looks for US businesses that hold the video you describe, reviews rights and consents, and delivers under a license that defines the records, allowed uses, term and delivery. Personal details are removed or replaced before delivery, with the method recorded and a sample checked. Describe your footage and caption requirements at sourcex.si/buyers.
Sources
- Krishna, Hata, Ren, Fei-Fei, Niebles (arXiv; ICCV 2017), "Dense-Captioning Events in Videos" (2017). https://arxiv.org/abs/1705.00754?context=cs
- arXiv (1903.01945), "MS-TCN: Multi-Stage Temporal Convolutional Network for Action Segmentation" (2019). https://arxiv.org/pdf/1903.01945
- arXiv (2209.13064), "EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations" (2022). https://arxiv.org/pdf/2209.13064
- arXiv (2603.06865), "Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation" (2026). https://arxiv.org/pdf/2603.06865
- Snap Research (GitHub; CVPR 2024), "Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers" (2024). https://github.com/snap-research/Panda-70M
- Northcutt, Athalye, Mueller (arXiv; NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Pushkarna, Zaldivar, Kjartansson (Google Research; FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.