فهم الفيديو
فهم ما يجري في مقطع فيديو
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
ما الذي تعنيه هذه القدرة
Takes a video, optionally with a question, and outputs a description or answer about content, events and timeline — what happened after three minutes, or whether a movement was performed correctly. It adds a time dimension to image understanding and must distinguish events with similar frames but different order. Input is usually sampled into frames and compressed before entering the model.
كيف تُنفَّذ تقنيًا
Frames are sampled on a fixed or adaptive schedule, a vision encoder turns each into tokens, and frame order with timestamps enters a multimodal Transformer whose language side produces text. Long videos are managed with hierarchical summarisation, keyframe selection or memory mechanisms to cap the token budget. Precise localisation is handled by a dedicated temporal grounding head or by having the model output a time interval.
منتجات تمثيلية
5Gemini
2023نموذج عام متعدد الوسائط بطبيعته، مصمم لسياقات طويلة جدًا
GPT-4o
2024نموذج عام متعدد الوسائط بطبيعته، يجمع النص والصورة والصوت في مدخل واحد
Qwen
2023عائلة بأوزان مفتوحة تغطي أحجامًا متعددة، مع نسخ متعددة الوسائط
Doubao
2023نموذج المحادثة العام وتطبيق بايت دانس
Hunyuan
2023عائلة نماذج تينسنت العامة، مع نسخ بأوزان مفتوحة
المؤسسات ذات الصلة
الاستخدامات الشائعة
- Event search in surveillance and security footage
- Technique review in sport and fitness
- Overviews of lectures and meetings
- Moderation and locating offending segments
كيف يُقاس مدى جودتها
- Video QA accuracy
- Share answered correctly on annotated questions
- Temporal grounding mIoU
- Mean IoU between predicted and true time intervals
- Video captioning CIDEr
- Semantic match of captions against references
الحدود والصعوبات
- Frame sampling discards mid-video detail, and one-off events are the easiest to miss
- Fine temporal ordering — who did what first — is vague and often reversed
- Judgement is unstable when audio and picture conflict, as dubbed audio can make the model waver
المفاهيم الكامنة وراءها
التوليد متعدد الوسائط
نموذج واحد يتعلّم الكلام والرسم والحركة، بل ونمذجة العالم ثلاثي الأبعاد
آلية الانتباه
يستطيع كل موضع أن ينظر مباشرة إلى جميع المواضع الأخرى ويوزّع الانتباه حسب الصلة
معمارية Transformer
يستبدل النقل كلمةً بكلمة بغرفة يتحدث فيها الجميع معاً، فتصبح التبعيات البعيدة على مسافة خطوة واحدة