Compreensão de imagens e VQA
Olhar uma imagem e responder perguntas livres
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE ESTA CAPACIDADE SIGNIFICA
Takes an image, optionally with a text question, and outputs a natural-language description or answer — what the person is doing, where the scene might be. Unlike OCR it targets meaning rather than characters, and unlike classification it answers open-ended questions instead of choosing from a fixed label set.
Como é feita tecnicamente
The mainstream is a multimodal large model: a vision encoder splits the image into patches and encodes them as visual tokens, which are fed with text tokens into one Transformer so the language side produces the answer. Alignment training uses image-caption pairs with contrastive learning, and later instruction tuning teaches question answering. Small text and fine detail are preserved by splitting the image into higher-resolution patches.
Produtos representativos
9GPT-4o
2024Um modelo geral nativamente multimodal: texto, imagem e áudio por uma única porta
Gemini
2023Um modelo geral nativamente multimodal, feito para contextos muito longos
Claude
2023Um modelo de conversa geral conhecido por contexto longo e alinhamento de segurança
Qwen
2023Uma família de pesos abertos com muitos tamanhos e versões multimodais
Doubao
2023O modelo de conversa geral e o aplicativo da ByteDance
Kimi
2023Um assistente de conversa chinês conhecido por lidar com contextos longos
GLM
2023Um modelo geral chinês que começou com pré-treinamento de preenchimento de lacunas autorregressivo
Gemini
2024Uma entrada de chat que reúne busca, escritório e um modelo multimodal
ChatGPT
2022A janela de chat que colocou um grande modelo de linguagem nas mãos de todos
Organizações relacionadas
Usos típicos
- Accessibility and image narration
- Product and listing description from photos
- Assisted reading of industrial and medical images
- Photo management and content search
Como avaliar se funciona bem
- VQA accuracy
- Share answered correctly on annotated question sets
- Captioning CIDEr / SPICE
- Semantic match between captions and references
- Hallucination rate
- Share of captions mentioning objects not in the image
Limites e dificuldades
- Exact counting and spatial relations are unreliable; miscounting people and flipping left-right are common
- Small text, chart readings and dense tables in the image are frequently misread
- The model fills gaps with common sense, describing what should be there as if it were seen
Conceitos por trás
Geração multimodal
Um único modelo que aprende a falar, desenhar, mover-se e até modelar o mundo 3D
Mecanismo de atenção
Cada posição pode olhar diretamente para todas as outras e distribuir atenção conforme a relevância
Arquitetura Transformer
Substituir o revezamento palavra a palavra por uma sala onde todos falam ao mesmo tempo, deixando dependências distantes a um salto