Hiểu video
Hiểu điều gì diễn ra trong video
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
NĂNG LỰC NÀY NGHĨA LÀ GÌ
Takes a video, optionally with a question, and outputs a description or answer about content, events and timeline — what happened after three minutes, or whether a movement was performed correctly. It adds a time dimension to image understanding and must distinguish events with similar frames but different order. Input is usually sampled into frames and compressed before entering the model.
Làm ra sao về mặt kỹ thuật
Frames are sampled on a fixed or adaptive schedule, a vision encoder turns each into tokens, and frame order with timestamps enters a multimodal Transformer whose language side produces text. Long videos are managed with hierarchical summarisation, keyframe selection or memory mechanisms to cap the token budget. Precise localisation is handled by a dedicated temporal grounding head or by having the model output a time interval.
Sản phẩm tiêu biểu
5Gemini
2023Mô hình đa phương thức gốc, xử lý ngữ cảnh siêu dài
GPT-4o
2024Mô hình đa phương thức gốc, xử lý văn bản, hình ảnh và âm thanh qua một cửa vào
Qwen
2023Dòng trọng số mở với nhiều kích cỡ và bản đa phương thức
Doubao
2023Mô hình hội thoại phổ thông và ứng dụng của ByteDance
Hunyuan
2023Dòng mô hình phổ thông của Tencent, có bản trọng số mở
Tổ chức liên quan
Cách dùng tiêu biểu
- Event search in surveillance and security footage
- Technique review in sport and fitness
- Overviews of lectures and meetings
- Moderation and locating offending segments
Đánh giá nó tốt hay không thế nào
- Video QA accuracy
- Share answered correctly on annotated questions
- Temporal grounding mIoU
- Mean IoU between predicted and true time intervals
- Video captioning CIDEr
- Semantic match of captions against references
Ranh giới và điểm khó
- Frame sampling discards mid-video detail, and one-off events are the easiest to miss
- Fine temporal ordering — who did what first — is vague and often reversed
- Judgement is unstable when audio and picture conflict, as dubbed audio can make the model waver
Các khái niệm đằng sau
Sinh đa phương thức
Một mô hình học nói, vẽ, chuyển động — thậm chí mô hình hóa thế giới ba chiều
Cơ chế chú ý (Attention)
Mọi vị trí đều có thể nhìn thẳng vào mọi vị trí khác và phân bổ chú ý theo mức liên quan
Kiến trúc Transformer
Thay cách truyền từng từ bằng một phòng họp nơi mọi từ cùng lên tiếng, để phụ thuộc xa chỉ còn cách một bước