تخطٍّ إلى المحتوى
أطلس الذكاء الاصطناعي

فهم الصور والأسئلة البصرية

النظر إلى صورة والإجابة عن أسئلة مفتوحة

الفهم البصريمبتدئ #17
إدخالصورةنصنص

يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.

ما الذي تعنيه هذه القدرة

Takes an image, optionally with a text question, and outputs a natural-language description or answer — what the person is doing, where the scene might be. Unlike OCR it targets meaning rather than characters, and unlike classification it answers open-ended questions instead of choosing from a fixed label set.

كيف تُنفَّذ تقنيًا

The mainstream is a multimodal large model: a vision encoder splits the image into patches and encodes them as visual tokens, which are fed with text tokens into one Transformer so the language side produces the answer. Alignment training uses image-caption pairs with contrastive learning, and later instruction tuning teaches question answering. Small text and fine detail are preserved by splitting the image into higher-resolution patches.

منتجات تمثيلية

9

GPT-4o

2024
OpenAI

نموذج عام متعدد الوسائط بطبيعته، يجمع النص والصورة والصوت في مدخل واحد

نموذج مغلق
نصصورةصوتنصصوت

Gemini

2023
Google DeepMind

نموذج عام متعدد الوسائط بطبيعته، مصمم لسياقات طويلة جدًا

نموذج مغلق
نصصورةصوتفيديونص

Claude

2023
Anthropic

نموذج محادثة عام يشتهر بسياقه الطويل ومواءمته الأمنية

نموذج مغلق
نصصورةنص

Qwen

2023
Alibaba (Qwen)

عائلة بأوزان مفتوحة تغطي أحجامًا متعددة، مع نسخ متعددة الوسائط

نموذج أوزان مفتوحة
نصصورةنص

Doubao

2023
ByteDance (Seed)

نموذج المحادثة العام وتطبيق بايت دانس

نموذج مغلق
نصصورةنص

Kimi

2023
Moonshot AI

مساعد محادثة صيني يشتهر بمعالجة السياقات الطويلة

نموذج مغلق
نصنص

GLM

2023
Zhipu AI

نموذج عام صيني بدأ بالتدريب المسبق لملء الفراغات ذاتيًا

نموذج مغلق
نصنص

Gemini

2024
Google DeepMind

مدخل محادثة واحد يجمع البحث والتطبيقات المكتبية ونموذجًا متعدد الوسائط

تطبيق مجاني جزئيًا
نصصورةصوتنصصورةصوت

ChatGPT

2022
OpenAI

نافذة محادثة جعلت نموذج اللغة الكبير في متناول الجميع

تطبيق مجاني جزئيًا
نصصورةصوتنصصورةصوت

المؤسسات ذات الصلة

الاستخدامات الشائعة

  • Accessibility and image narration
  • Product and listing description from photos
  • Assisted reading of industrial and medical images
  • Photo management and content search

كيف يُقاس مدى جودتها

VQA accuracy
Share answered correctly on annotated question sets
Captioning CIDEr / SPICE
Semantic match between captions and references
Hallucination rate
Share of captions mentioning objects not in the image

الحدود والصعوبات

  • Exact counting and spatial relations are unreliable; miscounting people and flipping left-right are common
  • Small text, chart readings and dense tables in the image are frequently misread
  • The model fills gaps with common sense, describing what should be there as if it were seen

المفاهيم الكامنة وراءها