Skip to content
AI Atlas

Video Understanding

Understand what happens in a video

VideoIntermediate #25
inVideoTextText

WHAT THIS CAPABILITY MEANS

Takes a video, optionally with a question, and outputs a description or answer about content, events and timeline — what happened after three minutes, or whether a movement was performed correctly. It adds a time dimension to image understanding and must distinguish events with similar frames but different order. Input is usually sampled into frames and compressed before entering the model.

How it is done

Frames are sampled on a fixed or adaptive schedule, a vision encoder turns each into tokens, and frame order with timestamps enters a multimodal Transformer whose language side produces text. Long videos are managed with hierarchical summarisation, keyframe selection or memory mechanisms to cap the token budget. Precise localisation is handled by a dedicated temporal grounding head or by having the model output a time interval.

Representative products

5

Organizations involved

Typical uses

  • Event search in surveillance and security footage
  • Technique review in sport and fitness
  • Overviews of lectures and meetings
  • Moderation and locating offending segments

How it is evaluated

Video QA accuracy
Share answered correctly on annotated questions
Temporal grounding mIoU
Mean IoU between predicted and true time intervals
Video captioning CIDEr
Semantic match of captions against references

Limits and hard parts

  • Frame sampling discards mid-video detail, and one-off events are the easiest to miss
  • Fine temporal ordering — who did what first — is vague and often reversed
  • Judgement is unstable when audio and picture conflict, as dubbed audio can make the model waver

Concepts behind it