كشف الأجسام
من «ما الذي في الصورة» إلى «ماذا وأين وكم العدد»
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
التعريف
Object detection answers not only “what is in the image” but “where each object is,” typically returning a set of bounding boxes, each with a class label and a confidence score. It advances classification from one label per image to one label plus one location per object.
الحدس المباشر
Classification says “this is a living room”; detection says “this is a living room, the sofa is bottom-left, the TV mid-top, and the cat is on the windowsill.” The former needs to get one thing right; the latter must simultaneously get the count, the classes and the positions right within one image — the difficulty jumps from recognition to search.
The detection pipeline: the backbone yields features, the head predicts densely or via proposals, and NMS deduplicates into labelled boxes
The mAP progression of detectors on COCO: from two-stage Faster R-CNN to one-stage YOLO models and end-to-end DETR
طريقة العمل
- 01
Three steps: classification, localisation, detection
Classification outputs one label; localisation regresses one box under the assumption of a single object; detection must handle an unknown number of instances of varying size that occlude one another.
- 02
Two-stage: propose first, classify each proposal
The R-CNN family first uses a Region Proposal Network (RPN) to generate thousands of candidate boxes, then classifies each and regresses its coordinates. Accurate but slow, and dependent on proposal quality.
- 03
One-stage: predict classes and boxes in a single pass
YOLO and SSD predict densely on the feature map directly, turning detection into a regression problem. Fast, though early versions were weak on small objects and crowded scenes — largely fixed in later iterations.
- 04
Anchors → anchor-free → end-to-end
From predefined anchors, to anchor-free methods such as FCOS and CenterNet, to DETR, which emits boxes directly via set prediction with Hungarian matching — the detection head keeps getting simpler.
- 05
Post-processing: IoU and NMS
Intersection-over-union (IoU) measures how much two boxes overlap; non-maximum suppression (NMS) removes duplicate boxes for the same object; mAP aggregates the result over several IoU thresholds.
A typical detection head: one feature map branches into a classification branch and a box-regression branch, later merged into “boxes + classes”
The speed-accuracy trade-off: one-stage methods are generally faster, two-stage ones more accurate, while modern models push both directions at once (hover for coordinates)
مجالات الاستخدام
- Autonomous driving: detecting vehicles, pedestrians and traffic signs
- Security and retail: people counting and shelf inventory auditing
- Medical imaging: locating lesions and cells
- Remote sensing: detecting aircraft, ships and buildings
مفاهيم خاطئة شائعة
- High mAP does not mean deployable. Misses and false alarms rarely cost the same, so inspect the precision-recall curve in context rather than trusting one number.
- NMS is a hyperparameter. A threshold too low deletes neighbouring objects such as two people in a crowd; too high keeps duplicate boxes.
- Small objects remain a weak spot: downsampling erases them, while high resolution brings a compute cost.
مصطلحات أساسية
- Bounding box
- A rectangle represented as (x, y, w, h) or corner points
- IoU
- The ratio of the intersection to the union of two boxes
- NMS
- Non-maximum suppression, removing duplicate boxes
- mAP
- Mean average precision across classes and IoU thresholds