Image Understanding & VQA
Look at an image and answer free-form questions
WHAT THIS CAPABILITY MEANS
Takes an image, optionally with a text question, and outputs a natural-language description or answer — what the person is doing, where the scene might be. Unlike OCR it targets meaning rather than characters, and unlike classification it answers open-ended questions instead of choosing from a fixed label set.
How it is done
The mainstream is a multimodal large model: a vision encoder splits the image into patches and encodes them as visual tokens, which are fed with text tokens into one Transformer so the language side produces the answer. Alignment training uses image-caption pairs with contrastive learning, and later instruction tuning teaches question answering. Small text and fine detail are preserved by splitting the image into higher-resolution patches.
Representative products
9GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Gemini
2023A natively multimodal general model built for very long context
Claude
2023A general chat model known for long context and safety alignment
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
Doubao
2023ByteDance’s general chat model and application
Kimi
2023A Chinese chat assistant known for long-context handling
GLM
2023A Chinese general model that began with autoregressive blank-infilling pretraining
Gemini
2024One chat entry point that gathers search, office apps and a multimodal model
ChatGPT
2022The chat window that put a large language model in everyone’s hands
Organizations involved
Typical uses
- Accessibility and image narration
- Product and listing description from photos
- Assisted reading of industrial and medical images
- Photo management and content search
How it is evaluated
- VQA accuracy
- Share answered correctly on annotated question sets
- Captioning CIDEr / SPICE
- Semantic match between captions and references
- Hallucination rate
- Share of captions mentioning objects not in the image
Limits and hard parts
- Exact counting and spatial relations are unreliable; miscounting people and flipping left-right are common
- Small text, chart readings and dense tables in the image are frequently misread
- The model fills gaps with common sense, describing what should be there as if it were seen
Concepts behind it
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away