Videoanalyse
Verstehen, was in einem Video geschieht
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes a video, optionally with a question, and outputs a description or answer about content, events and timeline — what happened after three minutes, or whether a movement was performed correctly. It adds a time dimension to image understanding and must distinguish events with similar frames but different order. Input is usually sampled into frames and compressed before entering the model.
Wie sie technisch umgesetzt wird
Frames are sampled on a fixed or adaptive schedule, a vision encoder turns each into tokens, and frame order with timestamps enters a multimodal Transformer whose language side produces text. Long videos are managed with hierarchical summarisation, keyframe selection or memory mechanisms to cap the token budget. Precise localisation is handled by a dedicated temporal grounding head or by having the model output a time interval.
Repräsentative Produkte
5Gemini
2023Ein von Grund auf multimodales Allzweckmodell für sehr langen Kontext
GPT-4o
2024Ein von Grund auf multimodales Allzweckmodell: Text, Bild und Audio über einen Zugang
Qwen
2023Eine Open-Weights-Familie über viele Größen hinweg, mit multimodalen Versionen
Doubao
2023ByteDances allgemeines Dialogmodell und App
Hunyuan
2023Tencent Generalmodell-Familie mit Open-Weights-Versionen
Beteiligte Organisationen
Typische Verwendungen
- Event search in surveillance and security footage
- Technique review in sport and fitness
- Overviews of lectures and meetings
- Moderation and locating offending segments
Wie sie bewertet wird
- Video QA accuracy
- Share answered correctly on annotated questions
- Temporal grounding mIoU
- Mean IoU between predicted and true time intervals
- Video captioning CIDEr
- Semantic match of captions against references
Grenzen und schwierige Punkte
- Frame sampling discards mid-video detail, and one-off events are the easiest to miss
- Fine temporal ordering — who did what first — is vague and often reversed
- Judgement is unstable when audio and picture conflict, as dubbed audio can make the model waver
Konzepte dahinter
Multimodale Generierung
Ein Modell, das sprechen, zeichnen, sich bewegen — und sogar die 3D-Welt modellieren lernt
Attention-Mechanismus
Jede Position kann direkt auf alle anderen blicken und ihre Aufmerksamkeit nach Relevanz verteilen
Transformer-Architektur
Statt Wort-für-Wort-Stafette ein Raum, in dem alle zugleich sprechen — und weite Abhängigkeiten sind nur einen Schritt entfernt