Object Detection
Draw a box around each object and name it
WHAT THIS CAPABILITY MEANS
Takes an image and outputs a set of bounding boxes, each with a class label and a confidence score. Unlike classification it is not content with one label for the whole image but must say what is where; unlike segmentation it gives rectangles rather than pixel-accurate outlines. Several instances of the same class can be detected at once.
How it is done
The mainstream is single-stage detection: anchors or centre points are densely predicted on a feature map and one forward pass regresses both box location and class at once, the YOLO line and SSD being the classic examples, fast enough for real time. Two-stage detectors first propose regions then classify each, slightly more accurate but slower. The loss optimises localisation and classification together, and non-maximum suppression merges overlapping boxes at the end.
Representative products
4Gemini
2023A natively multimodal general model built for very long context
GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
Waymo Driver
2009A self-driving system running driverless commercial robotaxis
Organizations involved
Typical uses
- Vehicle and pedestrian perception for driving
- Video surveillance and perimeter alerts
- Shelf and inventory counting in retail
- Object surveys in remote-sensing imagery
How it is evaluated
- mAP
- Mean of per-class average precision, usually at an IoU threshold
- IoU
- Intersection over union between predicted and true boxes, the hit criterion
- Inference speed (FPS)
- As important as accuracy in real-time settings
Limits and hard parts
- Recall drops sharply for occluded, truncated and very small objects
- In crowded scenes non-maximum suppression wrongly removes neighbours; two people side by side may become one
- Objects outside the trained classes are either missed or forced into the nearest known class
Concepts behind it
Object Detection
From “what is in the image” to “what, where, and how many”
Convolutional Neural Networks
Replacing full connections with “look locally, reuse the same filter everywhere” — the idea that made image recognition work
Image Representation
To a machine, a photo is nothing but stacked grids of numbers