Comprensión de vídeo
Entender qué ocurre en un vídeo
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ SIGNIFICA ESTA CAPACIDAD
Takes a video, optionally with a question, and outputs a description or answer about content, events and timeline — what happened after three minutes, or whether a movement was performed correctly. It adds a time dimension to image understanding and must distinguish events with similar frames but different order. Input is usually sampled into frames and compressed before entering the model.
Cómo se consigue técnicamente
Frames are sampled on a fixed or adaptive schedule, a vision encoder turns each into tokens, and frame order with timestamps enters a multimodal Transformer whose language side produces text. Long videos are managed with hierarchical summarisation, keyframe selection or memory mechanisms to cap the token budget. Precise localisation is handled by a dedicated temporal grounding head or by having the model output a time interval.
Productos representativos
5Gemini
2023Un modelo general nativamente multimodal, pensado para contextos muy largos
GPT-4o
2024Un modelo general nativamente multimodal: texto, imagen y audio por una misma puerta
Qwen
2023Una familia de pesos abiertos con muchos tamaños y versiones multimodales
Doubao
2023El modelo de chat general y la aplicación de ByteDance
Hunyuan
2023La familia de modelos generales de Tencent, con versiones de pesos abiertos
Organizaciones relacionadas
Usos típicos
- Event search in surveillance and security footage
- Technique review in sport and fitness
- Overviews of lectures and meetings
- Moderation and locating offending segments
Cómo se evalúa
- Video QA accuracy
- Share answered correctly on annotated questions
- Temporal grounding mIoU
- Mean IoU between predicted and true time intervals
- Video captioning CIDEr
- Semantic match of captions against references
Límites y dificultades
- Frame sampling discards mid-video detail, and one-off events are the easiest to miss
- Fine temporal ordering — who did what first — is vague and often reversed
- Judgement is unstable when audio and picture conflict, as dubbed audio can make the model waver
Conceptos detrás
Generación multimodal
Un mismo modelo que aprende a hablar, dibujar, moverse e incluso modelar el mundo 3D
Mecanismo de atención
Cada posición puede mirar directamente a todas las demás y repartir su atención según la relevancia
Arquitectura Transformer
Sustituir el relé palabra a palabra por una sala donde todos hablan a la vez, para que las dependencias lejanas estén a un salto