Bounding Box Annotation
Drawing rectangular boxes around target objects in images.
What this means for trainers
Consistency in how you handle partial occlusion and object edges matters more than raw speed — reviewers can spot inconsistent boxing conventions across a batch quickly.
Bounding box annotation is the foundational computer vision labeling task: drawing the tightest rectangle that fully contains a target object in an image, then tagging it with a class label. It is the entry point into object detection datasets and typically the first technique annotators learn before moving to more precise tasks like segmentation.
Quality here comes down to tightness and consistency. Boxes that are too loose include background noise that confuses the model during training, while boxes that clip part of the object teach it an incomplete shape. Partially occluded objects are a common edge case that guidelines need to address explicitly, since annotators need a clear rule for whether to box only the visible portion or estimate the object's full extent.
More complex spatial annotation tasks, such as 3D point cloud annotation and video event annotation, build on the same core skill of precisely bounding an object in space or time, which is why bounding boxes are usually where computer vision annotation training starts.
Related guides
Put this into practice
Browse open data annotation and evaluation roles from Mercor, Micro1, Outlier, and more.
Browse AI training jobs