Video Understanding
Understand what happens in a video
WHAT THIS CAPABILITY MEANS
Takes a video, optionally with a question, and outputs a description or answer about content, events and timeline — what happened after three minutes, or whether a movement was performed correctly. It adds a time dimension to image understanding and must distinguish events with similar frames but different order. Input is usually sampled into frames and compressed before entering the model.
How it is done
Frames are sampled on a fixed or adaptive schedule, a vision encoder turns each into tokens, and frame order with timestamps enters a multimodal Transformer whose language side produces text. Long videos are managed with hierarchical summarisation, keyframe selection or memory mechanisms to cap the token budget. Precise localisation is handled by a dedicated temporal grounding head or by having the model output a time interval.
Representative products
5Gemini
2023A natively multimodal general model built for very long context
GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
Doubao
2023ByteDance’s general chat model and application
Hunyuan
2023Tencent’s general model family, with open-weight versions
Organizations involved
Typical uses
- Event search in surveillance and security footage
- Technique review in sport and fitness
- Overviews of lectures and meetings
- Moderation and locating offending segments
How it is evaluated
- Video QA accuracy
- Share answered correctly on annotated questions
- Temporal grounding mIoU
- Mean IoU between predicted and true time intervals
- Video captioning CIDEr
- Semantic match of captions against references
Limits and hard parts
- Frame sampling discards mid-video detail, and one-off events are the easiest to miss
- Fine temporal ordering — who did what first — is vague and often reversed
- Judgement is unstable when audio and picture conflict, as dubbed audio can make the model waver
Concepts behind it
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away