Оценка позы и ключевых точек
Найти суставы и восстановить скелет
Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ
Takes an image or video and outputs coordinates for a set of keypoints — shoulders, elbows, wrists, hips, knees, ankles — which join into a skeleton. It adds geometric structure on top of detection and returns coordinate sequences rather than classes. It works both on a single person and on many people, grouping points per individual.
Как это устроено
Two paradigms dominate: top-down detects each person first and regresses keypoints inside the box, accurate but slower as the crowd grows; bottom-up predicts all joints over the image at once and assembles points into individuals using part-affinity fields, with speed largely independent of headcount. Heatmap regression was long the standard, and direct coordinate regression with Transformer backbones has since matured.
Примеры продуктов
4Gemini
2023Универсальная модель с нативной мультимодальностью, рассчитанная на очень длинный контекст
Qwen
2023Семейство с открытыми весами, охватывающее разные размеры и мультимодальные версии
Optimus
2022Проект человекоподобного робота, переиспользующий восприятие автопилота
Figure 02
2024Человекоподобный робот второго поколения под управлением модели «зрение–язык–действие»
Связанные организации
Типичное применение
- Fitness and sports motion analysis
- Human–computer interaction and gesture control
- Motion capture and animation driving
- Hand and face tracking
Как её оценивают
- PCK
- Share of keypoints falling within a radius of the truth
- OKS / keypoint mAP
- Detection average precision weighted by joint visibility
- MPJPE
- Mean joint position error in 3D pose, in millimetres
Границы и трудности
- Occlusion and crops drop keypoints, yet the model still fills in a plausible but wrong location
- Extreme poses — handstands, curled-up bodies — fall outside training and amplify error
- With mutual occlusion, limbs are stitched onto the wrong person
Концепции в основе
Обнаружение объектов
От «что на изображении» к «что, где и сколько»
Свёрточные нейронные сети
Замена полных связей на «смотреть локально и переиспользовать один и тот же фильтр везде» — идея, сделавшая распознавание изображений рабочим
Цифровое представление изображения
Для машины фотография — лишь набор наложенных сеток чисел