Estimación de pose y puntos clave
Localizar articulaciones y reconstruir el esqueleto
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ SIGNIFICA ESTA CAPACIDAD
Takes an image or video and outputs coordinates for a set of keypoints — shoulders, elbows, wrists, hips, knees, ankles — which join into a skeleton. It adds geometric structure on top of detection and returns coordinate sequences rather than classes. It works both on a single person and on many people, grouping points per individual.
Cómo se consigue técnicamente
Two paradigms dominate: top-down detects each person first and regresses keypoints inside the box, accurate but slower as the crowd grows; bottom-up predicts all joints over the image at once and assembles points into individuals using part-affinity fields, with speed largely independent of headcount. Heatmap regression was long the standard, and direct coordinate regression with Transformer backbones has since matured.
Productos representativos
4Gemini
2023Un modelo general nativamente multimodal, pensado para contextos muy largos
Qwen
2023Una familia de pesos abiertos con muchos tamaños y versiones multimodales
Optimus
2022Un proyecto de robot humanoide que reutiliza la percepción del autoconducción
Figure 02
2024Un humanoide de segunda generación movido por un modelo de visión-lenguaje-acción
Organizaciones relacionadas
Usos típicos
- Fitness and sports motion analysis
- Human–computer interaction and gesture control
- Motion capture and animation driving
- Hand and face tracking
Cómo se evalúa
- PCK
- Share of keypoints falling within a radius of the truth
- OKS / keypoint mAP
- Detection average precision weighted by joint visibility
- MPJPE
- Mean joint position error in 3D pose, in millimetres
Límites y dificultades
- Occlusion and crops drop keypoints, yet the model still fills in a plausible but wrong location
- Extreme poses — handstands, curled-up bodies — fall outside training and amplify error
- With mutual occlusion, limbs are stitched onto the wrong person
Conceptos detrás
Detección de objetos
De “qué hay en la imagen” a “qué, dónde y cuántos”
Redes neuronales convolucionales
Sustituir las conexiones completas por “mirar en local y reutilizar la misma regla en todas partes”: la idea que hizo funcionar el reconocimiento de imágenes
Representación digital de imágenes
Para una máquina, una foto no es más que cuadrículas de números superpuestas