Comprensión de imágenes y VQA
Mirar una imagen y responder preguntas libres
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ SIGNIFICA ESTA CAPACIDAD
Takes an image, optionally with a text question, and outputs a natural-language description or answer — what the person is doing, where the scene might be. Unlike OCR it targets meaning rather than characters, and unlike classification it answers open-ended questions instead of choosing from a fixed label set.
Cómo se consigue técnicamente
The mainstream is a multimodal large model: a vision encoder splits the image into patches and encodes them as visual tokens, which are fed with text tokens into one Transformer so the language side produces the answer. Alignment training uses image-caption pairs with contrastive learning, and later instruction tuning teaches question answering. Small text and fine detail are preserved by splitting the image into higher-resolution patches.
Productos representativos
9GPT-4o
2024Un modelo general nativamente multimodal: texto, imagen y audio por una misma puerta
Gemini
2023Un modelo general nativamente multimodal, pensado para contextos muy largos
Claude
2023Un modelo de chat general conocido por su contexto largo y su alineación de seguridad
Qwen
2023Una familia de pesos abiertos con muchos tamaños y versiones multimodales
Doubao
2023El modelo de chat general y la aplicación de ByteDance
Kimi
2023Un asistente de chat chino conocido por manejar contextos largos
GLM
2023Un modelo general chino que comenzó con preentrenamiento de relleno de espacios autorregresivo
Gemini
2024Una entrada de chat que reúne búsqueda, ofimática y un modelo multimodal
ChatGPT
2022La ventana de chat que puso un gran modelo de lenguaje en manos de todos
Organizaciones relacionadas
Usos típicos
- Accessibility and image narration
- Product and listing description from photos
- Assisted reading of industrial and medical images
- Photo management and content search
Cómo se evalúa
- VQA accuracy
- Share answered correctly on annotated question sets
- Captioning CIDEr / SPICE
- Semantic match between captions and references
- Hallucination rate
- Share of captions mentioning objects not in the image
Límites y dificultades
- Exact counting and spatial relations are unreliable; miscounting people and flipping left-right are common
- Small text, chart readings and dense tables in the image are frequently misread
- The model fills gaps with common sense, describing what should be there as if it were seen
Conceptos detrás
Generación multimodal
Un mismo modelo que aprende a hablar, dibujar, moverse e incluso modelar el mundo 3D
Mecanismo de atención
Cada posición puede mirar directamente a todas las demás y repartir su atención según la relevancia
Arquitectura Transformer
Sustituir el relé palabra a palabra por una sala donde todos hablan a la vez, para que las dependencias lejanas estén a un salto