मुद्रा और कीपॉइंट अनुमान
जोड़ों को खोजकर कंकाल पुनर्प्राप्त करना
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्षमता क्या है
Takes an image or video and outputs coordinates for a set of keypoints — shoulders, elbows, wrists, hips, knees, ankles — which join into a skeleton. It adds geometric structure on top of detection and returns coordinate sequences rather than classes. It works both on a single person and on many people, grouping points per individual.
तकनीकी रूप से कैसे
Two paradigms dominate: top-down detects each person first and regresses keypoints inside the box, accurate but slower as the crowd grows; bottom-up predicts all joints over the image at once and assembles points into individuals using part-affinity fields, with speed largely independent of headcount. Heatmap regression was long the standard, and direct coordinate regression with Transformer backbones has since matured.
प्रतिनिधि उत्पाद
4Gemini
2023मूल रूप से बहुविध, अति-लंबे संदर्भ के लिए बना सामान्य मॉडल
Qwen
2023अनेक आकारों और बहुविध संस्करणों वाला ओपन-वेट परिवार
Optimus
2022स्व-चालन की दृष्टि का पुनरुपयोग करने वाला मानव-सदृश रोबोट प्रोजेक्ट
Figure 02
2024दृष्टि-भाषा-क्रिया मॉडल से चलने वाला दूसरी पीढ़ी का मानव-सदृश रोबोट
संबंधित संस्थान
सामान्य उपयोग
- Fitness and sports motion analysis
- Human–computer interaction and gesture control
- Motion capture and animation driving
- Hand and face tracking
इसका मूल्यांकन कैसे होता है
- PCK
- Share of keypoints falling within a radius of the truth
- OKS / keypoint mAP
- Detection average precision weighted by joint visibility
- MPJPE
- Mean joint position error in 3D pose, in millimetres
सीमाएँ और कठिनाइयाँ
- Occlusion and crops drop keypoints, yet the model still fills in a plausible but wrong location
- Extreme poses — handstands, curled-up bodies — fall outside training and amplify error
- With mutual occlusion, limbs are stitched onto the wrong person
इसके पीछे की अवधारणाएँ
वस्तु संसूचन
“चित्र में क्या है” से “क्या, कहाँ और कितने” तक
संवलनीय तंत्रिका जाल
पूर्ण संयोजन की जगह «स्थानीय रूप से देखो, हर जगह वही पैमाना दोहराओ» — यही विचार छवि पहचान को वास्तव में कारगर बनाया
छवि का डिजिटल निरूपण
मशीन के लिए तस्वीर संख्याओं के परतदार जाल के अलावा कुछ नहीं