मुख्य सामग्री पर जाएँ

वस्तु पहचान

हर वस्तु पर बॉक्स बनाकर नाम बताना

दृष्टि बोधप्रारंभिक #12
इनपुटइमेजटेबल

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

यह क्षमता क्या है

Takes an image and outputs a set of bounding boxes, each with a class label and a confidence score. Unlike classification it is not content with one label for the whole image but must say what is where; unlike segmentation it gives rectangles rather than pixel-accurate outlines. Several instances of the same class can be detected at once.

तकनीकी रूप से कैसे

The mainstream is single-stage detection: anchors or centre points are densely predicted on a feature map and one forward pass regresses both box location and class at once, the YOLO line and SSD being the classic examples, fast enough for real time. Two-stage detectors first propose regions then classify each, slightly more accurate but slower. The loss optimises localisation and classification together, and non-maximum suppression merges overlapping boxes at the end.

प्रतिनिधि उत्पाद

4

संबंधित संस्थान

सामान्य उपयोग

  • Vehicle and pedestrian perception for driving
  • Video surveillance and perimeter alerts
  • Shelf and inventory counting in retail
  • Object surveys in remote-sensing imagery

इसका मूल्यांकन कैसे होता है

mAP
Mean of per-class average precision, usually at an IoU threshold
IoU
Intersection over union between predicted and true boxes, the hit criterion
Inference speed (FPS)
As important as accuracy in real-time settings

सीमाएँ और कठिनाइयाँ

  • Recall drops sharply for occluded, truncated and very small objects
  • In crowded scenes non-maximum suppression wrongly removes neighbours; two people side by side may become one
  • Objects outside the trained classes are either missed or forced into the nearest known class

इसके पीछे की अवधारणाएँ