動画理解
動画の中で何が起きたかを理解する
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes a video, optionally with a question, and outputs a description or answer about content, events and timeline — what happened after three minutes, or whether a movement was performed correctly. It adds a time dimension to image understanding and must distinguish events with similar frames but different order. Input is usually sampled into frames and compressed before entering the model.
技術的にどう実現するか
Frames are sampled on a fixed or adaptive schedule, a vision encoder turns each into tokens, and frame order with timestamps enters a multimodal Transformer whose language side produces text. Long videos are managed with hierarchical summarisation, keyframe selection or memory mechanisms to cap the token budget. Precise localisation is handled by a dedicated temporal grounding head or by having the model output a time interval.
代表的な製品
5Gemini
2023ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル
GPT-4o
2024ネイティブにマルチモーダルな汎用モデル。テキスト・画像・音声をひとつの入口で扱う
Qwen
2023多様な規模とマルチモーダル版を備えたオープンウェイトのモデル群
Doubao
2023バイトダンスの汎用対話モデルとアプリ
Hunyuan
2023開放ウェイト版を含むテンセントの汎用モデル群
関連する組織
代表的な用途
- Event search in surveillance and security footage
- Technique review in sport and fitness
- Overviews of lectures and meetings
- Moderation and locating offending segments
どう評価するか
- Video QA accuracy
- Share answered correctly on annotated questions
- Temporal grounding mIoU
- Mean IoU between predicted and true time intervals
- Video captioning CIDEr
- Semantic match of captions against references
限界と難しさ
- Frame sampling discards mid-video detail, and one-off events are the easiest to miss
- Fine temporal ordering — who did what first — is vague and often reversed
- Judgement is unstable when audio and picture conflict, as dubbed audio can make the model waver