فهم الصور والأسئلة البصرية
النظر إلى صورة والإجابة عن أسئلة مفتوحة
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
ما الذي تعنيه هذه القدرة
Takes an image, optionally with a text question, and outputs a natural-language description or answer — what the person is doing, where the scene might be. Unlike OCR it targets meaning rather than characters, and unlike classification it answers open-ended questions instead of choosing from a fixed label set.
كيف تُنفَّذ تقنيًا
The mainstream is a multimodal large model: a vision encoder splits the image into patches and encodes them as visual tokens, which are fed with text tokens into one Transformer so the language side produces the answer. Alignment training uses image-caption pairs with contrastive learning, and later instruction tuning teaches question answering. Small text and fine detail are preserved by splitting the image into higher-resolution patches.
منتجات تمثيلية
9GPT-4o
2024نموذج عام متعدد الوسائط بطبيعته، يجمع النص والصورة والصوت في مدخل واحد
Gemini
2023نموذج عام متعدد الوسائط بطبيعته، مصمم لسياقات طويلة جدًا
Claude
2023نموذج محادثة عام يشتهر بسياقه الطويل ومواءمته الأمنية
Qwen
2023عائلة بأوزان مفتوحة تغطي أحجامًا متعددة، مع نسخ متعددة الوسائط
Doubao
2023نموذج المحادثة العام وتطبيق بايت دانس
Kimi
2023مساعد محادثة صيني يشتهر بمعالجة السياقات الطويلة
GLM
2023نموذج عام صيني بدأ بالتدريب المسبق لملء الفراغات ذاتيًا
Gemini
2024مدخل محادثة واحد يجمع البحث والتطبيقات المكتبية ونموذجًا متعدد الوسائط
ChatGPT
2022نافذة محادثة جعلت نموذج اللغة الكبير في متناول الجميع
المؤسسات ذات الصلة
الاستخدامات الشائعة
- Accessibility and image narration
- Product and listing description from photos
- Assisted reading of industrial and medical images
- Photo management and content search
كيف يُقاس مدى جودتها
- VQA accuracy
- Share answered correctly on annotated question sets
- Captioning CIDEr / SPICE
- Semantic match between captions and references
- Hallucination rate
- Share of captions mentioning objects not in the image
الحدود والصعوبات
- Exact counting and spatial relations are unreliable; miscounting people and flipping left-right are common
- Small text, chart readings and dense tables in the image are frequently misread
- The model fills gaps with common sense, describing what should be there as if it were seen
المفاهيم الكامنة وراءها
التوليد متعدد الوسائط
نموذج واحد يتعلّم الكلام والرسم والحركة، بل ونمذجة العالم ثلاثي الأبعاد
آلية الانتباه
يستطيع كل موضع أن ينظر مباشرة إلى جميع المواضع الأخرى ويوزّع الانتباه حسب الصلة
معمارية Transformer
يستبدل النقل كلمةً بكلمة بغرفة يتحدث فيها الجميع معاً، فتصبح التبعيات البعيدة على مسافة خطوة واحدة