Pose & Keypoint Estimation
Locate joints and recover a skeleton
WHAT THIS CAPABILITY MEANS
Takes an image or video and outputs coordinates for a set of keypoints — shoulders, elbows, wrists, hips, knees, ankles — which join into a skeleton. It adds geometric structure on top of detection and returns coordinate sequences rather than classes. It works both on a single person and on many people, grouping points per individual.
How it is done
Two paradigms dominate: top-down detects each person first and regresses keypoints inside the box, accurate but slower as the crowd grows; bottom-up predicts all joints over the image at once and assembles points into individuals using part-affinity fields, with speed largely independent of headcount. Heatmap regression was long the standard, and direct coordinate regression with Transformer backbones has since matured.
Representative products
4Gemini
2023A natively multimodal general model built for very long context
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
Optimus
2022A humanoid robot project that reuses Tesla’s self-driving perception
Figure 02
2024A second-generation humanoid driven by a vision-language-action model
Organizations involved
Typical uses
- Fitness and sports motion analysis
- Human–computer interaction and gesture control
- Motion capture and animation driving
- Hand and face tracking
How it is evaluated
- PCK
- Share of keypoints falling within a radius of the truth
- OKS / keypoint mAP
- Detection average precision weighted by joint visibility
- MPJPE
- Mean joint position error in 3D pose, in millimetres
Limits and hard parts
- Occlusion and crops drop keypoints, yet the model still fills in a plausible but wrong location
- Extreme poses — handstands, curled-up bodies — fall outside training and amplify error
- With mutual occlusion, limbs are stitched onto the wrong person
Concepts behind it
Object Detection
From “what is in the image” to “what, where, and how many”
Convolutional Neural Networks
Replacing full connections with “look locally, reuse the same filter everywhere” — the idea that made image recognition work
Image Representation
To a machine, a photo is nothing but stacked grids of numbers