Text-to-Video
Turn a sentence into a short video clip
WHAT THIS CAPABILITY MEANS
Takes a text description and outputs a video with a time dimension, where frames must stay coherent with one another. It adds a hard constraint over text-to-image: temporal consistency, so objects cannot deform or teleport between neighbouring frames. Unlike image-to-video there is no starting frame; everything is determined by the prompt.
How it is done
The mainstream extends diffusion from two dimensions to three: a diffusion Transformer models space–time patches jointly so attention across frames constrains motion; another line uses a latent video autoencoder with spatio-temporal attention. Variable length and resolution are managed by generating and stitching in chunks. Synchronising audio and controlling camera language are recent extensions.
Representative products
8Sora
2024Generates coherent video up to about a minute long from a description
Veo
2024Generates 1080p video clips with coherent shots
Gen-3 Alpha
2024A highly controllable text-to-video model for film and advertising
Kling
2024A short-video model for both text-to-video and image-to-video
Hailuo
2024A short-video model focused on instruction following and camera language
Seedance
2024A video-generation model aimed at multi-shot storytelling
Dream Machine
2024Generates short videos with motion from text or an image
Synthesia
2019Type text, get a talking-avatar video
Organizations involved
Typical uses
- Fast spots for ads and short films
- Storyboards and animated pre-visualisation
- Short-form social media assets
- Shot drafts for games and virtual production
How it is evaluated
- FVD
- Distance between generated and real video distributions; lower is better
- Motion consistency
- Whether position and shape stay continuous across frames
- Human preference
- Pairwise comparison of visual quality and prompt adherence
Limits and hard parts
- Over long windows physics and causality break: objects vanish, merge or change identity
- Complex interactions and hand motion distort most, especially where several objects touch
- Clip length and resolution are compute-bound; beyond tens of seconds stitching shows seams
Concepts behind it
Diffusion Models
Learn a thousand tiny denoising steps, and you can build an image from pure noise
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world
Latent Diffusion & Conditional Control
Run diffusion not over pixels, but inside a compressed semantic space