画像理解と視覚的質問応答
画像を見て、それに関する自由な質問に答える
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes an image, optionally with a text question, and outputs a natural-language description or answer — what the person is doing, where the scene might be. Unlike OCR it targets meaning rather than characters, and unlike classification it answers open-ended questions instead of choosing from a fixed label set.
技術的にどう実現するか
The mainstream is a multimodal large model: a vision encoder splits the image into patches and encodes them as visual tokens, which are fed with text tokens into one Transformer so the language side produces the answer. Alignment training uses image-caption pairs with contrastive learning, and later instruction tuning teaches question answering. Small text and fine detail are preserved by splitting the image into higher-resolution patches.
代表的な製品
9GPT-4o
2024ネイティブにマルチモーダルな汎用モデル。テキスト・画像・音声をひとつの入口で扱う
Gemini
2023ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル
Claude
2023長い文脈と安全性の調整で知られる汎用対話モデル
Qwen
2023多様な規模とマルチモーダル版を備えたオープンウェイトのモデル群
Doubao
2023バイトダンスの汎用対話モデルとアプリ
Kimi
2023長い文脈の処理を得意とする中国語の対話アシスタント
GLM
2023自己回帰空白穴埋め事前学習から始まった中国語の汎用モデル
Gemini
2024検索・オフィス・マルチモーダルモデルを一つの対話入口に集約
ChatGPT
2022大規模言語モデルを誰もが使える対話画面にした
関連する組織
代表的な用途
- Accessibility and image narration
- Product and listing description from photos
- Assisted reading of industrial and medical images
- Photo management and content search
どう評価するか
- VQA accuracy
- Share answered correctly on annotated question sets
- Captioning CIDEr / SPICE
- Semantic match between captions and references
- Hallucination rate
- Share of captions mentioning objects not in the image
限界と難しさ
- Exact counting and spatial relations are unreliable; miscounting people and flipping left-right are common
- Small text, chart readings and dense tables in the image are frequently misread
- The model fills gaps with common sense, describing what should be there as if it were seen