テキストから動画生成
一文の説明から短い動画を生成する
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes a text description and outputs a video with a time dimension, where frames must stay coherent with one another. It adds a hard constraint over text-to-image: temporal consistency, so objects cannot deform or teleport between neighbouring frames. Unlike image-to-video there is no starting frame; everything is determined by the prompt.
技術的にどう実現するか
The mainstream extends diffusion from two dimensions to three: a diffusion Transformer models space–time patches jointly so attention across frames constrains motion; another line uses a latent video autoencoder with spatio-temporal attention. Variable length and resolution are managed by generating and stitching in chunks. Synchronising audio and controlling camera language are recent extensions.
代表的な製品
8Sora
2024一文の説明から最長約1分の一貫した動画を生成する
Veo
20241080pでショットの一貫した動画クリップを生成する
Gen-3 Alpha
2024映像・広告向けの制御性の高いテキスト動画生成モデル
Kling
2024テキストからも画像からも動画を生成できるショート動画モデル
Hailuo
2024指示追従とカメラワークを重視したショート動画モデル
Seedance
2024マルチショットの物語表現を狙った動画生成モデル
Dream Machine
2024テキストや画像から動きのある短い動画を生成する
Synthesia
2019テキストを入力するとアバターが話す動画ができる
関連する組織
代表的な用途
- Fast spots for ads and short films
- Storyboards and animated pre-visualisation
- Short-form social media assets
- Shot drafts for games and virtual production
どう評価するか
- FVD
- Distance between generated and real video distributions; lower is better
- Motion consistency
- Whether position and shape stay continuous across frames
- Human preference
- Pairwise comparison of visual quality and prompt adherence
限界と難しさ
- Over long windows physics and causality break: objects vanish, merge or change identity
- Complex interactions and hand motion distort most, especially where several objects touch
- Clip length and resolution are compute-bound; beyond tens of seconds stitching shows seams