Image data
Bounding Boxes vs Polygons vs Masks vs Keypoints: Which Annotation to Buy
Quick answer
Buy the cheapest annotation geometry that still encodes what your model must learn. Bounding boxes suit detection and counting when location is enough; polygons suit objects whose outline matters but where a few vertices capture it; pixel masks suit segmentation, area measurement and overlapping instances; keypoints suit pose, landmarks and part localization. Cost rises with geometric detail, but so does control over label noise. Specify the geometry, tolerance and export format in the purchase, then pilot before committing volume.
By SourceX Editorial · Updated
The model's output decides the annotation type
The annotation type you buy should match the output head you will train, not the richest label a vendor can produce. Encord frames the choice as what the model must learn: where an object is, what shape it has, which pixels belong to it, or how its parts relate [1]. A YOLO-style box detector uses a pixel mask only after it is collapsed to a box, while a Mask R-CNN or SAM-style fine-tune cannot learn boundaries from boxes alone without weak-supervision tricks.
Work backward from the metric your team will report. If the acceptance test is COCO-style box mAP, buying masks pays for precision you never score. If the downstream decision depends on defect area, crack length or canopy coverage, boxes will systematically mismeasure, and no amount of extra box volume fixes that.
| Model task | Minimum geometry | Typical metric | Where buyers overspend |
|---|---|---|---|
| Object detection, counting | Axis-aligned bounding box | box mAP@0.5, mAP@0.5:0.95 | Masks bought "for later" |
| Oriented objects (ships, shelf facings at angles) | Rotated box or 4-point polygon | rotated IoU mAP | Full polygons on rigid shapes |
| Instance segmentation | Polygon or mask per instance | mask AP | Pixel-perfect masks on soft edges |
| Semantic segmentation | Per-pixel class mask | mIoU | Instance IDs nobody uses |
| Pose, landmarks, part location | Keypoints with visibility flags | OKS-based AP, PCK | Masks for articulated subjects |
| Area or length measurement | Mask or polyline | Measurement error vs ground truth | Boxes as a cheap proxy |
Bounding boxes: cheapest per object, noisiest at the edges
Bounding boxes are the fastest geometry to draw and review, which is why most detection datasets use them. Dense retail benchmarks such as SKU-110K are labeled with boxes [7], since shelf detection needs location and count, not product outlines. For detection and tracking where objects are roughly rectangular, boxes are usually the right buy.
The trade-off is label noise inside the box. A study of annotation cost for computer-aided detection treats boxes as cheaper but noisier, lower-quality alternatives to contour annotations [2]. The noise is geometric: a box around a thin, diagonal or irregular object contains a large share of background pixels. That background becomes training signal, and for elongated objects such as cracks, cables, weld seams or surgical tools it can dominate the box.
Specify in the contract:
- Tightness rule: box edges touch the outermost visible pixel, with a stated pixel tolerance.
- Occlusion rule: box the visible extent only, or the inferred full extent (pick one and keep it consistent).
- Truncation and minimum-size rules: whether objects cut by the image border or smaller than N pixels are labeled.
- Crowd handling: whether dense groups get one crowd box or individual boxes.
Polygons: outline fidelity priced by vertex count
Polygons capture object shape at a cost that scales mostly with the number of vertices an annotator places. BasicAI publishes per-object price ranges by annotation type and notes that polygon cost depends on how complex the outline is, recommending a pilot before any estimate (vendor claim) [4]. Treat any per-object rate you see as a starting point; the real driver is your vertex density requirement and the share of complex objects in your images.
Polygons can also improve detection, not only segmentation. Roboflow reports an experiment in which a model trained on polygon labels reached higher mAP50 than the same setup trained on boxes (vendor-reported, single experiment) [3]. The plausible mechanism is that tighter geometry gives augmentation and box derivation better information, but verify on your own validation set before paying the premium.
Specify in the contract:
- Maximum boundary deviation in pixels, or a minimum vertex count per object class.
- How holes (a donut-shaped gasket, a window frame) are represented: separate inner polygons or a mask.
- Whether multi-part objects split by occlusion are one annotation with several polygons or separate instances.
Segmentation masks: pay for pixels only where pixels are scored
Pixel masks are the most expensive geometry and the only one that supports per-pixel metrics such as mIoU and direct area measurement. Semantic masks assign every pixel a class; instance masks separate each object; panoptic labels combine both. Masks are often produced today with model-assisted tools (click-to-segment, then human correction), so the cost of a mask depends heavily on how much correction the pre-label needs.
The research case for cheaper alternatives is real. Work on the cost-effectiveness of weak and noisy annotations for segmentation examines when boxes or coarse, noisy masks deliver comparable value per labeling dollar to precise masks [5]. A practical pattern is to buy precise masks for a validation and test set and a smaller training core, then use weaker labels for the bulk.
Mask failure modes to test in a pilot:
- Boundary halos: masks that systematically grow or shrink by a few pixels, which biases area estimates.
- Class bleed at contact points between adjacent instances.
- Inconsistent treatment of transparent, reflective or motion-blurred regions.
- RLE versus polygon export differences that change pixel counts after conversion.
Keypoints: structure, not extent
Keypoints encode where defined parts are, which suits human pose, hand and face landmarks, animal pose, vehicle corners and mechanical part alignment. Encord groups them with structural annotation for tasks where relationships between points matter more than object extent [1]. A keypoint schema is a contract in itself: the named points, their order, the skeleton edges and how invisible points are handled.
In COCO-format exports, keypoint annotations sit in a separate person_keypoints-style file, with each point stored as an x, y, visibility triplet and the keypoint names and skeleton defined on the category [6]. Ask suppliers to state whether visibility 1 (labeled but occluded) and visibility 2 (visible) are used consistently, because mixing them corrupts OKS-based evaluation. If subjects are people, keypoints on faces and bodies can raise biometric questions in some US states; route those to counsel before contracting.
Where cost actually comes from
Per-object price is a poor comparison unit across geometries, because object density, image resolution, class count and review depth vary more than the drawing time. The useful unit is cost per accepted annotation at your quality threshold, which includes rework. Vendor pricing pages publish ranges, not quotes, and recommend pilots for that reason [4].
The main cost drivers, roughly in order of impact:
- Objects per image: a shelf photo with 150 facings costs more than 150 single-object images to review.
- Geometry detail: vertex density for polygons, boundary tolerance for masks, point count for keypoints.
- Taxonomy size and ambiguity: more classes means more disagreement and adjudication.
- Review depth: single pass, consensus, or expert adjudication on a sample.
- Model-assisted pre-labeling: cuts drawing time but adds correction and audit effort.
For data you license rather than commission, the annotation already exists, so the question becomes whether its geometry and guidelines match your task. The trade-offs between licensing labeled sets and annotating raw images are covered in pre-labeled vs raw image datasets.
Annotation spec to attach to a purchase request
A written annotation spec turns the geometry choice into something a supplier can price and you can audit. Attach it to the request, alongside the taxonomy and guideline document described in writing image annotation guidelines and label taxonomies.
Illustrative example: invented to show structure; it does not describe an available dataset.
annotation_spec:
task: instance_segmentation # detection | instance_seg | semantic_seg | pose
geometry:
primary: polygon # bbox | rotated_bbox | polygon | mask | keypoints
boundary_tolerance_px: 2
holes: inner_polygons
occlusion: visible_extent_only
min_object_size_px: 16
classes:
- { id: 1, name: corrosion_patch }
- { id: 2, name: coating_failure }
keypoints: null # or { names: [...], skeleton: [[1,2],...], visibility: [0,1,2] }
export:
format: COCO_JSON # instances file; RLE allowed for crowd regions
bbox_convention: "[x, y, width, height], absolute pixels"
image_id_stable_across_versions: true
quality:
review: consensus_2_plus_adjudication
gold_set_share: 0.05
acceptance: "mask IoU >= 0.85 vs gold on sampled 300 objects"
pilot:
images: 200
report: [time_per_object, rework_rate, per_class_IoU]
Checking delivered labels against the geometry you paid for
Verify geometry on arrival with automated checks before any model training. In COCO JSON, the top-level sections are info, licenses, categories, images and annotations, and instance segmentations can be polygons or RLE [6]. That structure makes several checks cheap to script:
- Every annotation's bbox lies inside its image's width and height.
- Polygon vertex counts per class meet the spec minimum; flag degenerate polygons with fewer than three points.
- Mask area divided by box area per class falls in a plausible band; outliers reveal loose boxes or broken masks.
- Each keypoints array holds three values per defined point, num_keypoints equals the count of points with visibility above 0, and visibility values use only the agreed codes.
- A re-annotated gold sample hits the IoU or OKS threshold in the contract.
Label quality metrics and inter-annotator agreement are covered in depth in the training data quality assessment guide, and the broader image sourcing picture sits in the image datasets for computer vision hub. For definitions, see the data annotation glossary entry.
Sourcing labeled image data from operational records
Some of the most useful labeled images come from companies whose work already produces them: inspection photos with marked defects, claims images with adjuster notes, or field photos tied to work orders. SourceX sources operational datasets from US companies on request, including documents, engineering records and new recordings of hands-on work, and does not source scraped web content or generic CCTV or photos. Data is not held in stock, so a request does not guarantee a match, and every release is approved by the supplying company. Buyers can describe the image and label specification they need, and see expert annotations and labels for the licensing side.
Describe the labeled image data you need
SourceX looks for US businesses that hold the data you describe, reviews ownership and consents for each dataset, and delivers it under a license that defines the records, uses, term and delivery. Nothing is contracted until a supplier agrees, and pricing is set per deal. Start a buyer request at sourcex.si/buyers.
Sources
- Encord, "Bounding Box vs. Polygon vs. Segmentation vs. Keypoint: Which Annotation Type Fits Your Task?". https://encord.com/blog/choosing-image-annotation-type/
- arXiv, "Did You Get What You Paid For? Rethinking Annotation Cost of Deep Learning Based Computer Aided Detection" (2022). https://arxiv.org/pdf/2209.15314
- Roboflow, "Polygon vs. Bounding Box for Computer Vision Annotation". https://roboflow.com/blog/polygon-vs-bounding-box-computer-vision-annotation
- BasicAI, "Image Annotation Services Cost". https://www.basic.ai/blog-post/image-annotation-services-cost
- arXiv, "Cost-effectiveness of weak and noisy annotations for segmentation (arXiv:2312.10600v3)" (2023). https://arxiv.org/html/2312.10600v3
- CVAT.ai, "COCO (CVAT format documentation)". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
- Ultralytics, "SKU-110K Dataset". https://docs.ultralytics.com/datasets/detect/sku-110k
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.