Estimation de pose et de points clés
Localiser les articulations et reconstruire le squelette
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE DÉSIGNE CETTE CAPACITÉ
Takes an image or video and outputs coordinates for a set of keypoints — shoulders, elbows, wrists, hips, knees, ankles — which join into a skeleton. It adds geometric structure on top of detection and returns coordinate sequences rather than classes. It works both on a single person and on many people, grouping points per individual.
Comment c'est fait
Two paradigms dominate: top-down detects each person first and regresses keypoints inside the box, accurate but slower as the crowd grows; bottom-up predicts all joints over the image at once and assembles points into individuals using part-affinity fields, with speed largely independent of headcount. Heatmap regression was long the standard, and direct coordinate regression with Transformer backbones has since matured.
Produits représentatifs
4Gemini
2023Un modèle général nativement multimodal, conçu pour de très longs contextes
Qwen
2023Une famille à poids ouverts couvrant de nombreuses tailles, avec des versions multimodales
Optimus
2022Un projet de robot humanoïde qui réutilise la perception de la conduite autonome
Figure 02
2024Un humanoïde de deuxième génération piloté par un modèle vision-langage-action
Organisations concernées
Usages typiques
- Fitness and sports motion analysis
- Human–computer interaction and gesture control
- Motion capture and animation driving
- Hand and face tracking
Comment on l'évalue
- PCK
- Share of keypoints falling within a radius of the truth
- OKS / keypoint mAP
- Detection average precision weighted by joint visibility
- MPJPE
- Mean joint position error in 3D pose, in millimetres
Limites et points difficiles
- Occlusion and crops drop keypoints, yet the model still fills in a plausible but wrong location
- Extreme poses — handstands, curled-up bodies — fall outside training and amplify error
- With mutual occlusion, limbs are stitched onto the wrong person
Concepts sous-jacents
Détection d’objets
De « qu’y a-t-il dans l’image » à « quoi, où et combien »
Réseaux de neurones convolutifs
Remplacer les connexions denses par « regarder localement et réutiliser le même filtre partout » : l’idée qui a rendu la reconnaissance d’images enfin fonctionnelle
Représentation numérique de l’image
Pour une machine, une photo n’est qu’une pile de grilles de nombres