تخطٍّ إلى المحتوى
أطلس الذكاء الاصطناعي

النص إلى فيديو

تحويل جملة إلى مقطع فيديو

الفيديومبتدئ #23
إدخالنصفيديو

يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.

ما الذي تعنيه هذه القدرة

Takes a text description and outputs a video with a time dimension, where frames must stay coherent with one another. It adds a hard constraint over text-to-image: temporal consistency, so objects cannot deform or teleport between neighbouring frames. Unlike image-to-video there is no starting frame; everything is determined by the prompt.

كيف تُنفَّذ تقنيًا

The mainstream extends diffusion from two dimensions to three: a diffusion Transformer models space–time patches jointly so attention across frames constrains motion; another line uses a latent video autoencoder with spatio-temporal attention. Variable length and resolution are managed by generating and stitching in chunks. Synchronising audio and controlling camera language are recent extensions.

منتجات تمثيلية

8

المؤسسات ذات الصلة

الاستخدامات الشائعة

  • Fast spots for ads and short films
  • Storyboards and animated pre-visualisation
  • Short-form social media assets
  • Shot drafts for games and virtual production

كيف يُقاس مدى جودتها

FVD
Distance between generated and real video distributions; lower is better
Motion consistency
Whether position and shape stay continuous across frames
Human preference
Pairwise comparison of visual quality and prompt adherence

الحدود والصعوبات

  • Over long windows physics and causality break: objects vanish, merge or change identity
  • Complex interactions and hand motion distort most, especially where several objects touch
  • Clip length and resolution are compute-bound; beyond tens of seconds stitching shows seams

المفاهيم الكامنة وراءها