Posen- und Keypoint-Schätzung
Gelenke finden und ein Skelett rekonstruieren
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes an image or video and outputs coordinates for a set of keypoints — shoulders, elbows, wrists, hips, knees, ankles — which join into a skeleton. It adds geometric structure on top of detection and returns coordinate sequences rather than classes. It works both on a single person and on many people, grouping points per individual.
Wie sie technisch umgesetzt wird
Two paradigms dominate: top-down detects each person first and regresses keypoints inside the box, accurate but slower as the crowd grows; bottom-up predicts all joints over the image at once and assembles points into individuals using part-affinity fields, with speed largely independent of headcount. Heatmap regression was long the standard, and direct coordinate regression with Transformer backbones has since matured.
Repräsentative Produkte
4Gemini
2023Ein von Grund auf multimodales Allzweckmodell für sehr langen Kontext
Qwen
2023Eine Open-Weights-Familie über viele Größen hinweg, mit multimodalen Versionen
Optimus
2022Ein humanoider Roboter, der die Wahrnehmung des Autopiloten wiederverwendet
Figure 02
2024Ein humanoider Roboter der zweiten Generation, gesteuert von einem Vision-Sprache-Aktion-Modell
Beteiligte Organisationen
Typische Verwendungen
- Fitness and sports motion analysis
- Human–computer interaction and gesture control
- Motion capture and animation driving
- Hand and face tracking
Wie sie bewertet wird
- PCK
- Share of keypoints falling within a radius of the truth
- OKS / keypoint mAP
- Detection average precision weighted by joint visibility
- MPJPE
- Mean joint position error in 3D pose, in millimetres
Grenzen und schwierige Punkte
- Occlusion and crops drop keypoints, yet the model still fills in a plausible but wrong location
- Extreme poses — handstands, curled-up bodies — fall outside training and amplify error
- With mutual occlusion, limbs are stitched onto the wrong person
Konzepte dahinter
Objekterkennung
Von „was ist im Bild“ zu „was, wo und wie viele“
Faltungsnetze
Vollverbindungen durch „lokal schauen, überall denselben Filter wiederverwenden“ ersetzen – die Idee, die Bilderkennung erst brauchbar machte
Digitale Bilddarstellung
Für eine Maschine ist ein Foto nichts als gestapelte Zahlenraster