Detecção de objetos
Desenhar uma caixa por objeto e nomeá-lo
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE ESTA CAPACIDADE SIGNIFICA
Takes an image and outputs a set of bounding boxes, each with a class label and a confidence score. Unlike classification it is not content with one label for the whole image but must say what is where; unlike segmentation it gives rectangles rather than pixel-accurate outlines. Several instances of the same class can be detected at once.
Como é feita tecnicamente
The mainstream is single-stage detection: anchors or centre points are densely predicted on a feature map and one forward pass regresses both box location and class at once, the YOLO line and SSD being the classic examples, fast enough for real time. Two-stage detectors first propose regions then classify each, slightly more accurate but slower. The loss optimises localisation and classification together, and non-maximum suppression merges overlapping boxes at the end.
Produtos representativos
4Gemini
2023Um modelo geral nativamente multimodal, feito para contextos muito longos
GPT-4o
2024Um modelo geral nativamente multimodal: texto, imagem e áudio por uma única porta
Qwen
2023Uma família de pesos abertos com muitos tamanhos e versões multimodais
Waymo Driver
2009Um sistema de direção autônoma que opera robotáxis comerciais sem motorista de segurança
Organizações relacionadas
Usos típicos
- Vehicle and pedestrian perception for driving
- Video surveillance and perimeter alerts
- Shelf and inventory counting in retail
- Object surveys in remote-sensing imagery
Como avaliar se funciona bem
- mAP
- Mean of per-class average precision, usually at an IoU threshold
- IoU
- Intersection over union between predicted and true boxes, the hit criterion
- Inference speed (FPS)
- As important as accuracy in real-time settings
Limites e dificuldades
- Recall drops sharply for occluded, truncated and very small objects
- In crowded scenes non-maximum suppression wrongly removes neighbours; two people side by side may become one
- Objects outside the trained classes are either missed or forced into the nearest known class
Conceitos por trás
Detecção de objetos
De “o que há na imagem” para “o quê, onde e quantos”
Redes neurais convolucionais
Substituir as conexões densas por “olhar localmente e reutilizar o mesmo filtro em toda parte”: a ideia que tornou o reconhecimento de imagens funcional
Representação digital de imagens
Para uma máquina, uma foto não passa de grades de números sobrepostas