छवि समझ और विज़ुअल प्रश्नोत्तर
छवि देखकर उस पर खुले प्रश्नों का उत्तर देना
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्षमता क्या है
Takes an image, optionally with a text question, and outputs a natural-language description or answer — what the person is doing, where the scene might be. Unlike OCR it targets meaning rather than characters, and unlike classification it answers open-ended questions instead of choosing from a fixed label set.
तकनीकी रूप से कैसे
The mainstream is a multimodal large model: a vision encoder splits the image into patches and encodes them as visual tokens, which are fed with text tokens into one Transformer so the language side produces the answer. Alignment training uses image-caption pairs with contrastive learning, and later instruction tuning teaches question answering. Small text and fine detail are preserved by splitting the image into higher-resolution patches.
प्रतिनिधि उत्पाद
9GPT-4o
2024मूल रूप से बहुविध सामान्य मॉडल — पाठ, चित्र और ऑडियो एक ही द्वार से
Gemini
2023मूल रूप से बहुविध, अति-लंबे संदर्भ के लिए बना सामान्य मॉडल
Claude
2023लंबे संदर्भ और सुरक्षा-संरेखण के लिए जाना जाने वाला सामान्य संवाद मॉडल
Qwen
2023अनेक आकारों और बहुविध संस्करणों वाला ओपन-वेट परिवार
Doubao
2023बाइटडांस का सामान्य संवाद मॉडल और ऐप
Kimi
2023लंबे संदर्भ के लिए जाना जाने वाला चीनी संवाद सहायक
GLM
2023ऑटोरेग्रेसिव ब्लैंक-भरने वाले प्रीट्रेनिंग से शुरू हुआ चीनी सामान्य मॉडल
Gemini
2024खोज, ऑफ़िस और बहुविध मॉडल को एक चैट द्वार में समेटता है
ChatGPT
2022वह चैट विंडो जिसने बड़े भाषा मॉडल को हर किसी तक पहुँचाया
संबंधित संस्थान
सामान्य उपयोग
- Accessibility and image narration
- Product and listing description from photos
- Assisted reading of industrial and medical images
- Photo management and content search
इसका मूल्यांकन कैसे होता है
- VQA accuracy
- Share answered correctly on annotated question sets
- Captioning CIDEr / SPICE
- Semantic match between captions and references
- Hallucination rate
- Share of captions mentioning objects not in the image
सीमाएँ और कठिनाइयाँ
- Exact counting and spatial relations are unreliable; miscounting people and flipping left-right are common
- Small text, chart readings and dense tables in the image are frequently misread
- The model fills gaps with common sense, describing what should be there as if it were seen
इसके पीछे की अवधारणाएँ
बहु-मॉडल जनरेशन
एक ही मॉडल बोलना, चित्र बनाना, हिलना, और यहाँ तक कि 3D संसार का नमूना बनाना सीखता है
अटेंशन तंत्र
हर स्थान बाकी सभी स्थानों को सीधे देख सकता है और प्रासंगिकता के अनुसार ध्यान बाँट सकता है
Transformer आर्किटेक्चर
शब्द-दर-शब्द relay की जगह वह कक्ष जहाँ सब एक साथ बोलते हैं, जिससे दूर की निर्भरताएँ एक कदम पर आ जाती हैं