Compreensão de vídeo
Entender o que acontece em um vídeo
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE ESTA CAPACIDADE SIGNIFICA
Takes a video, optionally with a question, and outputs a description or answer about content, events and timeline — what happened after three minutes, or whether a movement was performed correctly. It adds a time dimension to image understanding and must distinguish events with similar frames but different order. Input is usually sampled into frames and compressed before entering the model.
Como é feita tecnicamente
Frames are sampled on a fixed or adaptive schedule, a vision encoder turns each into tokens, and frame order with timestamps enters a multimodal Transformer whose language side produces text. Long videos are managed with hierarchical summarisation, keyframe selection or memory mechanisms to cap the token budget. Precise localisation is handled by a dedicated temporal grounding head or by having the model output a time interval.
Produtos representativos
5Gemini
2023Um modelo geral nativamente multimodal, feito para contextos muito longos
GPT-4o
2024Um modelo geral nativamente multimodal: texto, imagem e áudio por uma única porta
Qwen
2023Uma família de pesos abertos com muitos tamanhos e versões multimodais
Doubao
2023O modelo de conversa geral e o aplicativo da ByteDance
Hunyuan
2023A família de modelos gerais da Tencent, com versões de pesos abertos
Organizações relacionadas
Usos típicos
- Event search in surveillance and security footage
- Technique review in sport and fitness
- Overviews of lectures and meetings
- Moderation and locating offending segments
Como avaliar se funciona bem
- Video QA accuracy
- Share answered correctly on annotated questions
- Temporal grounding mIoU
- Mean IoU between predicted and true time intervals
- Video captioning CIDEr
- Semantic match of captions against references
Limites e dificuldades
- Frame sampling discards mid-video detail, and one-off events are the easiest to miss
- Fine temporal ordering — who did what first — is vague and often reversed
- Judgement is unstable when audio and picture conflict, as dubbed audio can make the model waver
Conceitos por trás
Geração multimodal
Um único modelo que aprende a falar, desenhar, mover-se e até modelar o mundo 3D
Mecanismo de atenção
Cada posição pode olhar diretamente para todas as outras e distribuir atenção conforme a relevância
Arquitetura Transformer
Substituir o revezamento palavra a palavra por uma sala onde todos falam ao mesmo tempo, deixando dependências distantes a um salto