Compréhension de vidéo
Comprendre ce qui se passe dans une vidéo
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE DÉSIGNE CETTE CAPACITÉ
Takes a video, optionally with a question, and outputs a description or answer about content, events and timeline — what happened after three minutes, or whether a movement was performed correctly. It adds a time dimension to image understanding and must distinguish events with similar frames but different order. Input is usually sampled into frames and compressed before entering the model.
Comment c'est fait
Frames are sampled on a fixed or adaptive schedule, a vision encoder turns each into tokens, and frame order with timestamps enters a multimodal Transformer whose language side produces text. Long videos are managed with hierarchical summarisation, keyframe selection or memory mechanisms to cap the token budget. Precise localisation is handled by a dedicated temporal grounding head or by having the model output a time interval.
Produits représentatifs
5Gemini
2023Un modèle général nativement multimodal, conçu pour de très longs contextes
GPT-4o
2024Un modèle général nativement multimodal : texte, image et audio par une même entrée
Qwen
2023Une famille à poids ouverts couvrant de nombreuses tailles, avec des versions multimodales
Doubao
2023Le modèle de conversation général et l’application de ByteDance
Hunyuan
2023La famille de modèles généraux de Tencent, avec des versions à poids ouverts
Organisations concernées
Usages typiques
- Event search in surveillance and security footage
- Technique review in sport and fitness
- Overviews of lectures and meetings
- Moderation and locating offending segments
Comment on l'évalue
- Video QA accuracy
- Share answered correctly on annotated questions
- Temporal grounding mIoU
- Mean IoU between predicted and true time intervals
- Video captioning CIDEr
- Semantic match of captions against references
Limites et points difficiles
- Frame sampling discards mid-video detail, and one-off events are the easiest to miss
- Fine temporal ordering — who did what first — is vague and often reversed
- Judgement is unstable when audio and picture conflict, as dubbed audio can make the model waver
Concepts sous-jacents
Génération multimodale
Un seul modèle qui apprend à parler, dessiner, bouger — et même à modéliser le monde 3D
Mécanisme d’attention
Chaque position peut regarder directement toutes les autres et répartir dynamiquement son attention selon la pertinence
Architecture Transformer
Remplacer le relais mot à mot par une salle où tous parlent en même temps, pour que les dépendances lointaines soient à un saut