Image-to-Video
Set a still image in motion
WHAT THIS CAPABILITY MEANS
Takes an image, optionally with a motion description, and outputs a clip that starts from it. The first frame is usually tightly constrained to the input, and later frames extrapolate motion and camera movement from there. Unlike text-to-video it has a definite visual starting point, which raises the bar for preserving subject identity and appearance.
How it is done
A common approach encodes the input as a condition, either pinning the first frame of the diffusion process or injecting it as a reference, then generates the following frames; another route extends keyframes with an image model and interpolates in between. Motion magnitude and camera control are set by explicit strength parameters or trajectory conditions, while physical plausibility is learned implicitly from motion priors in the data.
Representative products
8Stable Video Diffusion
2023Turns a single still image into a short video with a diffusion model
Kling
2024A short-video model for both text-to-video and image-to-video
Hailuo
2024A short-video model focused on instruction following and camera language
Dream Machine
2024Generates short videos with motion from text or an image
Veo
2024Generates 1080p video clips with coherent shots
Gen-3 Alpha
2024A highly controllable text-to-video model for film and advertising
Seedance
2024A video-generation model aimed at multi-shot storytelling
Sora
2024Generates coherent video up to about a minute long from a description
Organizations involved
Typical uses
- Animating photos and memory clips
- Product showcases and animated ad visuals
- Motion pre-vis from key art and storyboards
- Shot drafts for games and film
How it is evaluated
- FVD
- Distribution gap between generated and real video
- First-frame fidelity
- Agreement between the first frame and the input
- Human preference
- Pairwise judgement of motion naturalness and subject retention
Limits and hard parts
- Identity drifts after the first frame, with faces and clothing changing first
- Large camera moves break background geometry: straight lines bend and buildings misalign
- Occlusion and parallax are often wrong, inverting front-back relations
Concepts behind it
Diffusion Models
Learn a thousand tiny denoising steps, and you can build an image from pure noise
Latent Diffusion & Conditional Control
Run diffusion not over pixels, but inside a compressed semantic space
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world