वस्तु पहचान
हर वस्तु पर बॉक्स बनाकर नाम बताना
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्षमता क्या है
Takes an image and outputs a set of bounding boxes, each with a class label and a confidence score. Unlike classification it is not content with one label for the whole image but must say what is where; unlike segmentation it gives rectangles rather than pixel-accurate outlines. Several instances of the same class can be detected at once.
तकनीकी रूप से कैसे
The mainstream is single-stage detection: anchors or centre points are densely predicted on a feature map and one forward pass regresses both box location and class at once, the YOLO line and SSD being the classic examples, fast enough for real time. Two-stage detectors first propose regions then classify each, slightly more accurate but slower. The loss optimises localisation and classification together, and non-maximum suppression merges overlapping boxes at the end.
प्रतिनिधि उत्पाद
4Gemini
2023मूल रूप से बहुविध, अति-लंबे संदर्भ के लिए बना सामान्य मॉडल
GPT-4o
2024मूल रूप से बहुविध सामान्य मॉडल — पाठ, चित्र और ऑडियो एक ही द्वार से
Qwen
2023अनेक आकारों और बहुविध संस्करणों वाला ओपन-वेट परिवार
Waymo Driver
2009बिना सुरक्षा-चालक के व्यावसायिक रोबोटैक्सी चलाने वाला स्व-चालन तंत्र
संबंधित संस्थान
सामान्य उपयोग
- Vehicle and pedestrian perception for driving
- Video surveillance and perimeter alerts
- Shelf and inventory counting in retail
- Object surveys in remote-sensing imagery
इसका मूल्यांकन कैसे होता है
- mAP
- Mean of per-class average precision, usually at an IoU threshold
- IoU
- Intersection over union between predicted and true boxes, the hit criterion
- Inference speed (FPS)
- As important as accuracy in real-time settings
सीमाएँ और कठिनाइयाँ
- Recall drops sharply for occluded, truncated and very small objects
- In crowded scenes non-maximum suppression wrongly removes neighbours; two people side by side may become one
- Objects outside the trained classes are either missed or forced into the nearest known class
इसके पीछे की अवधारणाएँ
वस्तु संसूचन
“चित्र में क्या है” से “क्या, कहाँ और कितने” तक
संवलनीय तंत्रिका जाल
पूर्ण संयोजन की जगह «स्थानीय रूप से देखो, हर जगह वही पैमाना दोहराओ» — यही विचार छवि पहचान को वास्तव में कारगर बनाया
छवि का डिजिटल निरूपण
मशीन के लिए तस्वीर संख्याओं के परतदार जाल के अलावा कुछ नहीं