Skip to content
AI Atlas

Object Detection

From “what is in the image” to “what, where, and how many”

05 Computer VisionIntermediateEntry 4 in this domain

DEFINITION

Object detection answers not only “what is in the image” but “where each object is,” typically returning a set of bounding boxes, each with a class label and a confidence score. It advances classification from one label per image to one label plus one location per object.

Intuition

Classification says “this is a living room”; detection says “this is a living room, the sofa is bottom-left, the TV mid-top, and the cat is on the windowsill.” The former needs to get one thing right; the latter must simultaneously get the count, the classes and the positions right within one image — the difficulty jumps from recognition to search.

Fig. 1

The detection pipeline: the backbone yields features, the head predicts densely or via proposals, and NMS deduplicates into labelled boxes

Fig. 2

The mAP progression of detectors on COCO: from two-stage Faster R-CNN to one-stage YOLO models and end-to-end DETR

How it works

  1. 01

    Three steps: classification, localisation, detection

    Classification outputs one label; localisation regresses one box under the assumption of a single object; detection must handle an unknown number of instances of varying size that occlude one another.

  2. 02

    Two-stage: propose first, classify each proposal

    The R-CNN family first uses a Region Proposal Network (RPN) to generate thousands of candidate boxes, then classifies each and regresses its coordinates. Accurate but slow, and dependent on proposal quality.

  3. 03

    One-stage: predict classes and boxes in a single pass

    YOLO and SSD predict densely on the feature map directly, turning detection into a regression problem. Fast, though early versions were weak on small objects and crowded scenes — largely fixed in later iterations.

  4. 04

    Anchors → anchor-free → end-to-end

    From predefined anchors, to anchor-free methods such as FCOS and CenterNet, to DETR, which emits boxes directly via set prediction with Hungarian matching — the detection head keeps getting simpler.

  5. 05

    Post-processing: IoU and NMS

    Intersection-over-union (IoU) measures how much two boxes overlap; non-maximum suppression (NMS) removes duplicate boxes for the same object; mAP aggregates the result over several IoU thresholds.

Fig. 3

A typical detection head: one feature map branches into a classification branch and a box-regression branch, later merged into “boxes + classes”

Fig. 4

The speed-accuracy trade-off: one-stage methods are generally faster, two-stage ones more accurate, while modern models push both directions at once (hover for coordinates)

Where it is used

  • Autonomous driving: detecting vehicles, pedestrians and traffic signs
  • Security and retail: people counting and shelf inventory auditing
  • Medical imaging: locating lesions and cells
  • Remote sensing: detecting aircraft, ships and buildings

Common misconceptions

  • High mAP does not mean deployable. Misses and false alarms rarely cost the same, so inspect the precision-recall curve in context rather than trusting one number.
  • NMS is a hyperparameter. A threshold too low deletes neighbouring objects such as two people in a crowd; too high keeps duplicate boxes.
  • Small objects remain a weak spot: downsampling erases them, while high resolution brings a compute cost.

Key terms

Bounding box
A rectangle represented as (x, y, w, h) or corner points
IoU
The ratio of the intersection to the union of two boxes
NMS
Non-maximum suppression, removing duplicate boxes
mAP
Mean average precision across classes and IoU thresholds

Further reading