객체 탐지
“무엇이 있는가”에서 “무엇이·어디에·몇 개 있는가”로
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
정의
Object detection answers not only “what is in the image” but “where each object is,” typically returning a set of bounding boxes, each with a class label and a confidence score. It advances classification from one label per image to one label plus one location per object.
직관적 이해
Classification says “this is a living room”; detection says “this is a living room, the sofa is bottom-left, the TV mid-top, and the cat is on the windowsill.” The former needs to get one thing right; the latter must simultaneously get the count, the classes and the positions right within one image — the difficulty jumps from recognition to search.
The detection pipeline: the backbone yields features, the head predicts densely or via proposals, and NMS deduplicates into labelled boxes
The mAP progression of detectors on COCO: from two-stage Faster R-CNN to one-stage YOLO models and end-to-end DETR
작동 원리
- 01
Three steps: classification, localisation, detection
Classification outputs one label; localisation regresses one box under the assumption of a single object; detection must handle an unknown number of instances of varying size that occlude one another.
- 02
Two-stage: propose first, classify each proposal
The R-CNN family first uses a Region Proposal Network (RPN) to generate thousands of candidate boxes, then classifies each and regresses its coordinates. Accurate but slow, and dependent on proposal quality.
- 03
One-stage: predict classes and boxes in a single pass
YOLO and SSD predict densely on the feature map directly, turning detection into a regression problem. Fast, though early versions were weak on small objects and crowded scenes — largely fixed in later iterations.
- 04
Anchors → anchor-free → end-to-end
From predefined anchors, to anchor-free methods such as FCOS and CenterNet, to DETR, which emits boxes directly via set prediction with Hungarian matching — the detection head keeps getting simpler.
- 05
Post-processing: IoU and NMS
Intersection-over-union (IoU) measures how much two boxes overlap; non-maximum suppression (NMS) removes duplicate boxes for the same object; mAP aggregates the result over several IoU thresholds.
A typical detection head: one feature map branches into a classification branch and a box-regression branch, later merged into “boxes + classes”
The speed-accuracy trade-off: one-stage methods are generally faster, two-stage ones more accurate, while modern models push both directions at once (hover for coordinates)
응용 분야
- Autonomous driving: detecting vehicles, pedestrians and traffic signs
- Security and retail: people counting and shelf inventory auditing
- Medical imaging: locating lesions and cells
- Remote sensing: detecting aircraft, ships and buildings
흔한 오해
- High mAP does not mean deployable. Misses and false alarms rarely cost the same, so inspect the precision-recall curve in context rather than trusting one number.
- NMS is a hyperparameter. A threshold too low deletes neighbouring objects such as two people in a crowd; too high keeps duplicate boxes.
- Small objects remain a weak spot: downsampling erases them, while high resolution brings a compute cost.
핵심 용어
- Bounding box
- A rectangle represented as (x, y, w, h) or corner points
- IoU
- The ratio of the intersection to the union of two boxes
- NMS
- Non-maximum suppression, removing duplicate boxes
- mAP
- Mean average precision across classes and IoU thresholds