Văn bản thành video
Biến một câu mô tả thành đoạn video ngắn
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
NĂNG LỰC NÀY NGHĨA LÀ GÌ
Takes a text description and outputs a video with a time dimension, where frames must stay coherent with one another. It adds a hard constraint over text-to-image: temporal consistency, so objects cannot deform or teleport between neighbouring frames. Unlike image-to-video there is no starting frame; everything is determined by the prompt.
Làm ra sao về mặt kỹ thuật
The mainstream extends diffusion from two dimensions to three: a diffusion Transformer models space–time patches jointly so attention across frames constrains motion; another line uses a latent video autoencoder with spatio-temporal attention. Variable length and resolution are managed by generating and stitching in chunks. Synchronising audio and controlling camera language are recent extensions.
Sản phẩm tiêu biểu
8Sora
2024Tạo video mạch lạc dài tới khoảng một phút từ một mô tả
Veo
2024Tạo clip video 1080p với các cảnh mạch lạc
Gen-3 Alpha
2024Mô hình văn bản thành video dễ điều khiển cho phim và quảng cáo
Kling
2024Mô hình video ngắn cho cả văn bản thành video và ảnh thành video
Hailuo
2024Mô hình video ngắn chú trọng tuân thủ chỉ dẫn và ngôn ngữ máy quay
Seedance
2024Mô hình tạo video hướng tới kể chuyện nhiều cảnh
Dream Machine
2024Tạo video ngắn có chuyển động từ văn bản hoặc ảnh
Synthesia
2019Nhập văn bản, nhận video có nhân vật ảo đang nói
Tổ chức liên quan
Cách dùng tiêu biểu
- Fast spots for ads and short films
- Storyboards and animated pre-visualisation
- Short-form social media assets
- Shot drafts for games and virtual production
Đánh giá nó tốt hay không thế nào
- FVD
- Distance between generated and real video distributions; lower is better
- Motion consistency
- Whether position and shape stay continuous across frames
- Human preference
- Pairwise comparison of visual quality and prompt adherence
Ranh giới và điểm khó
- Over long windows physics and causality break: objects vanish, merge or change identity
- Complex interactions and hand motion distort most, especially where several objects touch
- Clip length and resolution are compute-bound; beyond tens of seconds stitching shows seams
Các khái niệm đằng sau
Mô hình khuếch tán
Học ngàn bước khử nhiễu nhỏ, bạn tạo được ảnh từ nhiễu thuần
Sinh đa phương thức
Một mô hình học nói, vẽ, chuyển động — thậm chí mô hình hóa thế giới ba chiều
Khuếch tán trong không gian ẩn và điều khiển có điều kiện
Khuếch tán không trên điểm ảnh, mà trong một không gian ngữ nghĩa đã nén