Skip to content

Video data

Annotating Surgical Video: Phases, Steps, Instruments and Skill Ratings

Quick answer

A usable surgical phase recognition dataset is defined by its label specification more than its footage. Fix a procedure-specific phase and step ontology with written boundary rules, choose the cheapest instrument label that serves the model (presence, boxes, masks or tracks), score safety milestones such as the critical view of safety and technical skill on validated scales, and require clinically qualified annotators whose agreement is measured and reported per label type.

By SourceX Editorial · Updated

This page covers the annotation layer. For how hospitals share the underlying footage, consent and HIPAA questions, see surgical video datasets and how hospitals share footage; for the wider category, start at the video data hub.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Phase and step ontologies: start from a published per-procedure definition

Phase labels are only comparable when they follow a fixed, procedure-specific ontology, so begin with a published definition and document every deviation. The reference case is Cholec80, introduced with EndoNet: laparoscopic cholecystectomy videos annotated with seven phases (preparation, Calot triangle dissection, clipping and cutting, gallbladder dissection, gallbladder packaging, cleaning and coagulation, gallbladder retraction) plus tool presence [1]. Many cholecystectomy models are still benchmarked against those phase names, so a new dataset that renames or merges them loses comparability.

Separate three granularities and do not mix them in one track. Phases are coarse surgical goals lasting minutes; steps are sub-goals within a phase (for example, clip the cystic duct, then clip the cystic artery); actions or gestures are seconds-long motions such as "push needle through tissue." JIGSAWS is the canonical gesture-level example [7] [7]: a fixed gesture vocabulary applied to suturing, knot-tying and needle-passing on a bench-top model. The general mechanics of segment labels are covered in temporal action segmentation and localization labels.

For procedures without a published ontology (many robotic, orthopedic or endoscopic cases), write one with the clinical lead before any labeling. Each phase needs a definition, an onset cue, an offset cue, the allowed predecessors and successors, and a rule for revisits, because real cases return to earlier phases (bleeding control mid-dissection is common).

Boundary rules and tolerance: where phase labels actually disagree

Most disagreement between qualified annotators sits at transitions, not in the middle of phases, so the spec must say exactly which frame starts a phase and how much boundary slack evaluation will allow. Typical cues are instrument-driven ("clipping and cutting begins when the clip applier enters the field") or anatomy-driven ("dissection ends when the gallbladder is detached from the liver bed"). Write the cue, not the intent.

Define tolerances in seconds per transition and report boundary-aware metrics alongside frame accuracy. Frame accuracy hides short phases and over-segmentation; segment-level edit scores and F1 at overlap thresholds expose them. Also specify how to label idle time, camera out of the body, lens cleaning and smoke-obscured stretches, because those segments silently distort phase-duration statistics if annotators improvise.

Failure modes to check in a delivered set:

  • Off-by-a-cue labels where one annotator anchors on instrument entry and another on first tissue contact.
  • Phases forced into a linear order that the case did not follow.
  • Overlapping or gapped segments in the export (end of one phase not equal to start of the next).
  • Frame-rate drift: labels authored at one rate and video re-encoded at another, shifting every boundary.

Instrument labels: presence, boxes, masks and tracks are separate cost tiers

Choose the lowest instrument label tier that serves the task, because each step up multiplies annotation time and reviewer load. In rough order of effort: frame-level tool presence (multi-label binary per frame), bounding boxes, instance or semantic masks, then identity-consistent tracks across frames. Cholec80 shipped tool presence only [1]; CholecSeg8k later added pixel-level semantic masks on a subset of Cholec80 frames [2]; Endoscapes pairs boxes on 1,933 frames with masks on 422 frames from 50 videos [3].

That pattern, dense cheap labels across many videos and expensive masks on a sampled subset, is usually the right buying shape. Specify the instrument taxonomy explicitly: generic class ("grasper") versus model-specific class, how to label the shaft versus the jaw, whether partially visible or out-of-focus tools count, and whether tissue in the jaws belongs to the tool mask. For anatomy masks, list the structures (cystic duct, cystic artery, hepatocystic triangle, cystic plate) and the rule for occlusion by fat or blood.

Export formats matter for reuse. Ask for COCO-style JSON or per-frame PNG masks with a published class map, segment tables keyed to video ID and frame index, and the exact frame-extraction rate. Delivery conventions and schemas are covered in dataset delivery formats and transfer.

Critical view of safety: annotate criteria, not a single yes/no

Critical view of safety (CVS) labels should record each criterion separately, with multiple clinical raters, rather than one binary "CVS achieved" flag. Endoscapes is the public reference: 201 cholecystectomy videos with CVS assessed on 11,090 frames by three clinical experts, plus official splits and benchmarks for detection, segmentation and CVS prediction [3]. Keeping per-rater, per-criterion scores lets you study disagreement and choose an aggregation rule later instead of inheriting one.

In the spec, define the three criteria in the wording your clinical lead accepts, the visual evidence required for each, and whether raters score single frames or short clips. Decide up front whether the label is "visible in this frame" or "achieved by this point in the case," since those produce different training targets. Note the access model too: Endoscapes2023 is also published on PhysioNet as a restricted-access database [4], which is typical of surgical sets and affects whether you can use them for commercial training at all.

Skill ratings: use a validated scale and trained raters

Skill labels are only defensible when they come from a validated assessment scale applied by trained raters who are blind to the operator's identity and experience level. JIGSAWS rates trials with a modified OSATS global rating score, dropping categories that do not apply to controlled bench-top clips, alongside kinematics from the da Vinci system and stereo endoscopic video. That provenance is why it remains a benchmark, but it also covers only eight surgeons and three dry-lab tasks.

When specifying new skill data, name the scale (OSATS-style global ratings for open and bench tasks, procedure-specific scales where they exist), the number of raters per video, rater training and calibration videos, and blinding. Record the operator's training level separately as metadata, never as the label. Gesture and step labels drift as annotators tire or reinterpret the vocabulary, so budget a re-review pass for any fine-grained surgical label.

Annotator qualifications and layered QA

Phase, step, anatomy and CVS labels need clinically qualified annotators, with non-clinical labelers limited to tasks a clinician has made mechanical, such as tool boxes against a fixed taxonomy. A practical split: surgeons or senior residents write the ontology and adjudicate; trained clinical annotators label phases, steps and anatomy; trained generalists label instrument presence and boxes with clinical spot review.

Layer the QA. Large video segmentation benchmarks document qualification tasks before production, ongoing checks during labeling and review by additional annotators [5]. For surgical work, add gold videos with adjudicated labels seeded into each annotator's queue, double-labeling of a fixed share of videos, and adjudication of every disagreement above tolerance. For acceptance testing methods, see how to audit annotation quality.

Report agreement per label type with a metric that fits the data. Krippendorff's alpha handles multiple raters, missing ratings and nominal or ordinal values [6], which suits CVS criteria and ordinal skill scores; for phase boundaries, report boundary offset distributions in seconds rather than a single coefficient.

Label specification template

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldWhat to specifyExample entry
procedureProcedure and approachLaparoscopic cholecystectomy, elective
ontology_versionPhase/step list and sourceCholec80 phases v1 + 11 local steps, spec rev 3
boundary_ruleOnset/offset cue per transitionClipping starts at clip applier entry
tolerance_sAllowed boundary offset per transition2 s phases, 1 s steps
non_surgical_segmentsLabels for idle, out-of-body, lens cleaningOOB, IDLE, CLEAN
instrument_tierPresence, box, mask, trackPresence all frames; masks on 1 frame per 30 s
instrument_taxonomyClasses and part rules7 classes; jaw and shaft in one mask
cvs_scoringCriteria, unit, raters3 criteria, per frame, 3 raters, no forced consensus
skill_scaleScale, raters, blindingOSATS-style GRS, 2 blinded raters
annotator_rolesWho labels and adjudicates whatSurgeon adjudicator; clinical annotators; generalists for boxes
qaGold share, double-label share, metrics5% gold, 15% double-labeled, alpha and boundary offsets
exportFormats and keysCOCO JSON, PNG masks, CSV segments keyed to video_id + frame_idx at stated fps

A matching record for one phase segment might look like this:

{"video_id": "case_0417", "fps": 25, "track": "phase", "label": "clipping_and_cutting",
 "start_frame": 31250, "end_frame": 36874, "annotator_id": "A07", "adjudicated": true,
 "ontology_version": "chole-phases-r3", "boundary_cue": "clip_applier_entry"}

Where licensed surgical annotations come from

Public benchmarks such as Cholec80, Endoscapes and JIGSAWS are research resources with their own access terms, so commercial teams usually need footage and labels licensed for their actual use. Check each set's terms before training, and see license expert annotations and labels and the data annotation glossary entry for definitions.

SourceX sources operational datasets from US companies on request; categories are not inventory and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, health records require HIPAA de-identification by Safe Harbor or Expert Determination, and personal details are removed or replaced before delivery with the method recorded and a sample checked. No de-identification method is perfect, and video adds faces, voices and on-screen text, covered in rights layers in a video clip. Buyers can describe the footage and label specification on the SourceX buyer page.

Specify your surgical annotation request

SourceX looks for US businesses that hold the data you describe, assesses the data and licensing permissions, and agrees pricing and allowed uses in a license; nothing is contracted until a supplier agrees. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Start a buyer request at sourcex.si/buyers.

Frequently asked questions

Can I reuse Cholec80 phase names for a different procedure?

No. Phase ontologies are procedure-specific [1]; reuse only the structure (definitions, cues, tolerances) and have a clinical lead write phases for the new procedure.

Is frame-level tool presence enough for instrument segmentation?

Not for training a segmentation model, but it is a cheap first pass that lets you sample which frames are worth masking, as CholecSeg8k did on a Cholec80 subset [2].

Should CVS labels be forced to a consensus?

Keep each rater's per-criterion score and decide aggregation at training time; Endoscapes collected CVS ratings from three clinical experts per frame, which makes that kind of disagreement analysis possible [3].

Sources

  1. arXiv (Twinanda et al.), "EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic Videos" (2016). https://arxiv.org/pdf/1602.03012
  2. arXiv, "CholecSeg8k: A Semantic Segmentation Dataset for Laparoscopic Cholecystectomy Based on Cholec80" (2020). https://arxiv.org/abs/2012.12453v1
  3. arXiv (Murali, Alapatt, Mascagni et al.), "The Endoscapes Dataset for Surgical Scene Segmentation, Object Detection, and Critical View of Safety Assessment: Official Splits and Benchmark" (2023). https://arxiv.org/abs/2312.12429v1
  4. PhysioNet, "Endoscapes2023 (restricted-access database)" (2024). https://www.physionet.org/content/endoscapes-2023/
  5. arXiv, "EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations" (2022). https://arxiv.org/pdf/2209.13064
  6. University of Pennsylvania, Annenberg School for Communication (Krippendorff), "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  7. Gao et al., "The JHU-ISI Gesture and Skill Assessment Working Set (JIGSAWS)" (2014). https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.698.8028&rep=rep1&type=pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data