Phát hiện đối tượng
Vẽ khung quanh từng vật thể và gọi tên
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
NĂNG LỰC NÀY NGHĨA LÀ GÌ
Takes an image and outputs a set of bounding boxes, each with a class label and a confidence score. Unlike classification it is not content with one label for the whole image but must say what is where; unlike segmentation it gives rectangles rather than pixel-accurate outlines. Several instances of the same class can be detected at once.
Làm ra sao về mặt kỹ thuật
The mainstream is single-stage detection: anchors or centre points are densely predicted on a feature map and one forward pass regresses both box location and class at once, the YOLO line and SSD being the classic examples, fast enough for real time. Two-stage detectors first propose regions then classify each, slightly more accurate but slower. The loss optimises localisation and classification together, and non-maximum suppression merges overlapping boxes at the end.
Sản phẩm tiêu biểu
4Gemini
2023Mô hình đa phương thức gốc, xử lý ngữ cảnh siêu dài
GPT-4o
2024Mô hình đa phương thức gốc, xử lý văn bản, hình ảnh và âm thanh qua một cửa vào
Qwen
2023Dòng trọng số mở với nhiều kích cỡ và bản đa phương thức
Waymo Driver
2009Hệ thống tự lái vận hành taxi robot thương mại không có người giám sát
Tổ chức liên quan
Cách dùng tiêu biểu
- Vehicle and pedestrian perception for driving
- Video surveillance and perimeter alerts
- Shelf and inventory counting in retail
- Object surveys in remote-sensing imagery
Đánh giá nó tốt hay không thế nào
- mAP
- Mean of per-class average precision, usually at an IoU threshold
- IoU
- Intersection over union between predicted and true boxes, the hit criterion
- Inference speed (FPS)
- As important as accuracy in real-time settings
Ranh giới và điểm khó
- Recall drops sharply for occluded, truncated and very small objects
- In crowded scenes non-maximum suppression wrongly removes neighbours; two people side by side may become one
- Objects outside the trained classes are either missed or forced into the nearest known class
Các khái niệm đằng sau
Phát hiện đối tượng
Từ “trong ảnh có gì” đến “cái gì, ở đâu và bao nhiêu”
Mạng nơ-ron tích chập
Thay kết nối đầy đủ bằng “nhìn cục bộ, dùng lại cùng một thước ở mọi nơi” — ý tưởng đưa nhận dạng ảnh thành khả thi
Biểu diễn số của ảnh
Với máy móc, một bức ảnh chỉ là các lưới số xếp chồng