Estimativa de pose e pontos-chave
Localizar articulações e reconstruir o esqueleto
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE ESTA CAPACIDADE SIGNIFICA
Takes an image or video and outputs coordinates for a set of keypoints — shoulders, elbows, wrists, hips, knees, ankles — which join into a skeleton. It adds geometric structure on top of detection and returns coordinate sequences rather than classes. It works both on a single person and on many people, grouping points per individual.
Como é feita tecnicamente
Two paradigms dominate: top-down detects each person first and regresses keypoints inside the box, accurate but slower as the crowd grows; bottom-up predicts all joints over the image at once and assembles points into individuals using part-affinity fields, with speed largely independent of headcount. Heatmap regression was long the standard, and direct coordinate regression with Transformer backbones has since matured.
Produtos representativos
4Gemini
2023Um modelo geral nativamente multimodal, feito para contextos muito longos
Qwen
2023Uma família de pesos abertos com muitos tamanhos e versões multimodais
Optimus
2022Um projeto de robô humanoide que reaproveita a percepção da condução autônoma
Figure 02
2024Um humanoide de segunda geração guiado por um modelo de visão-linguagem-ação
Organizações relacionadas
Usos típicos
- Fitness and sports motion analysis
- Human–computer interaction and gesture control
- Motion capture and animation driving
- Hand and face tracking
Como avaliar se funciona bem
- PCK
- Share of keypoints falling within a radius of the truth
- OKS / keypoint mAP
- Detection average precision weighted by joint visibility
- MPJPE
- Mean joint position error in 3D pose, in millimetres
Limites e dificuldades
- Occlusion and crops drop keypoints, yet the model still fills in a plausible but wrong location
- Extreme poses — handstands, curled-up bodies — fall outside training and amplify error
- With mutual occlusion, limbs are stitched onto the wrong person
Conceitos por trás
Detecção de objetos
De “o que há na imagem” para “o quê, onde e quantos”
Redes neurais convolucionais
Substituir as conexões densas por “olhar localmente e reutilizar o mesmo filtro em toda parte”: a ideia que tornou o reconhecimento de imagens funcional
Representação digital de imagens
Para uma máquina, uma foto não passa de grades de números sobrepostas