Понимание видео
Понять, что происходит в видео
Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ
Takes a video, optionally with a question, and outputs a description or answer about content, events and timeline — what happened after three minutes, or whether a movement was performed correctly. It adds a time dimension to image understanding and must distinguish events with similar frames but different order. Input is usually sampled into frames and compressed before entering the model.
Как это устроено
Frames are sampled on a fixed or adaptive schedule, a vision encoder turns each into tokens, and frame order with timestamps enters a multimodal Transformer whose language side produces text. Long videos are managed with hierarchical summarisation, keyframe selection or memory mechanisms to cap the token budget. Precise localisation is handled by a dedicated temporal grounding head or by having the model output a time interval.
Примеры продуктов
5Gemini
2023Универсальная модель с нативной мультимодальностью, рассчитанная на очень длинный контекст
GPT-4o
2024Универсальная модель с нативной мультимодальностью: текст, изображение и звук через один вход
Qwen
2023Семейство с открытыми весами, охватывающее разные размеры и мультимодальные версии
Doubao
2023Универсальная диалоговая модель и приложение ByteDance
Hunyuan
2023Семейство универсальных моделей Tencent с открытыми весами
Связанные организации
Типичное применение
- Event search in surveillance and security footage
- Technique review in sport and fitness
- Overviews of lectures and meetings
- Moderation and locating offending segments
Как её оценивают
- Video QA accuracy
- Share answered correctly on annotated questions
- Temporal grounding mIoU
- Mean IoU between predicted and true time intervals
- Video captioning CIDEr
- Semantic match of captions against references
Границы и трудности
- Frame sampling discards mid-video detail, and one-off events are the easiest to miss
- Fine temporal ordering — who did what first — is vague and often reversed
- Judgement is unstable when audio and picture conflict, as dubbed audio can make the model waver
Концепции в основе
Мультимодальная генерация
Одна модель учится говорить, рисовать, двигаться — и даже моделировать трёхмерный мир
Механизм внимания
Каждая позиция может напрямую «видеть» все остальные и динамически распределять внимание по релевантности
Архитектура Transformer
Замена эстафеты по одному слову залом, где все говорят сразу, — и дальние зависимости оказываются в одном шаге