Detección de objetos
Dibujar una caja por objeto y nombrarlo
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ SIGNIFICA ESTA CAPACIDAD
Takes an image and outputs a set of bounding boxes, each with a class label and a confidence score. Unlike classification it is not content with one label for the whole image but must say what is where; unlike segmentation it gives rectangles rather than pixel-accurate outlines. Several instances of the same class can be detected at once.
Cómo se consigue técnicamente
The mainstream is single-stage detection: anchors or centre points are densely predicted on a feature map and one forward pass regresses both box location and class at once, the YOLO line and SSD being the classic examples, fast enough for real time. Two-stage detectors first propose regions then classify each, slightly more accurate but slower. The loss optimises localisation and classification together, and non-maximum suppression merges overlapping boxes at the end.
Productos representativos
4Gemini
2023Un modelo general nativamente multimodal, pensado para contextos muy largos
GPT-4o
2024Un modelo general nativamente multimodal: texto, imagen y audio por una misma puerta
Qwen
2023Una familia de pesos abiertos con muchos tamaños y versiones multimodales
Waymo Driver
2009Un sistema de conducción autónoma que opera robotaxis comerciales sin conductor de seguridad
Organizaciones relacionadas
Usos típicos
- Vehicle and pedestrian perception for driving
- Video surveillance and perimeter alerts
- Shelf and inventory counting in retail
- Object surveys in remote-sensing imagery
Cómo se evalúa
- mAP
- Mean of per-class average precision, usually at an IoU threshold
- IoU
- Intersection over union between predicted and true boxes, the hit criterion
- Inference speed (FPS)
- As important as accuracy in real-time settings
Límites y dificultades
- Recall drops sharply for occluded, truncated and very small objects
- In crowded scenes non-maximum suppression wrongly removes neighbours; two people side by side may become one
- Objects outside the trained classes are either missed or forced into the nearest known class
Conceptos detrás
Detección de objetos
De “qué hay en la imagen” a “qué, dónde y cuántos”
Redes neuronales convolucionales
Sustituir las conexiones completas por “mirar en local y reutilizar la misma regla en todas partes”: la idea que hizo funcionar el reconocimiento de imágenes
Representación digital de imágenes
Para una máquina, una foto no es más que cuadrículas de números superpuestas