टेक्स्ट-से-वीडियो
एक वाक्य से छोटा वीडियो बनाना
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्षमता क्या है
Takes a text description and outputs a video with a time dimension, where frames must stay coherent with one another. It adds a hard constraint over text-to-image: temporal consistency, so objects cannot deform or teleport between neighbouring frames. Unlike image-to-video there is no starting frame; everything is determined by the prompt.
तकनीकी रूप से कैसे
The mainstream extends diffusion from two dimensions to three: a diffusion Transformer models space–time patches jointly so attention across frames constrains motion; another line uses a latent video autoencoder with spatio-temporal attention. Variable length and resolution are managed by generating and stitching in chunks. Synchronising audio and controlling camera language are recent extensions.
प्रतिनिधि उत्पाद
8Sora
2024एक विवरण से लगभग एक मिनट तक का सुसंगत वीडियो बनाता है
Veo
2024सुसंगत शॉट के साथ 1080p वीडियो क्लिप बनाता है
Gen-3 Alpha
2024फ़िल्म और विज्ञापन के लिए अत्यधिक नियंत्रण योग्य टेक्स्ट-टू-वीडियो मॉडल
Kling
2024टेक्स्ट-टू-वीडियो और इमेज-टू-वीडियो दोनों के लिए शॉर्ट-वीडियो मॉडल
Hailuo
2024निर्देश-पालन और कैमरा भाषा पर केंद्रित शॉर्ट-वीडियो मॉडल
Seedance
2024मल्टी-शॉट कथा-कथन के लिए वीडियो जनरेटिंग मॉडल
Dream Machine
2024पाठ या चित्र से गति युक्त छोटे वीडियो बनाता है
Synthesia
2019टेक्स्ट लिखें, बोलते अवतार वाला वीडियो पाएँ
संबंधित संस्थान
सामान्य उपयोग
- Fast spots for ads and short films
- Storyboards and animated pre-visualisation
- Short-form social media assets
- Shot drafts for games and virtual production
इसका मूल्यांकन कैसे होता है
- FVD
- Distance between generated and real video distributions; lower is better
- Motion consistency
- Whether position and shape stay continuous across frames
- Human preference
- Pairwise comparison of visual quality and prompt adherence
सीमाएँ और कठिनाइयाँ
- Over long windows physics and causality break: objects vanish, merge or change identity
- Complex interactions and hand motion distort most, especially where several objects touch
- Clip length and resolution are compute-bound; beyond tens of seconds stitching shows seams
इसके पीछे की अवधारणाएँ
विसरण मॉडल
हज़ार छोटे शोर-हटाने के चरण सीखें, और शुद्ध शोर से चित्र बना सकेंगे
बहु-मॉडल जनरेशन
एक ही मॉडल बोलना, चित्र बनाना, हिलना, और यहाँ तक कि 3D संसार का नमूना बनाना सीखता है
अव्यक्त-समष्टि विसरण और सशर्त नियंत्रण
विसरण पिक्सेल पर नहीं, बल्कि संपीड़ित अर्थ-समष्टि के भीतर चलाएँ