Stable Video Diffusion
Turns a single still image into a short video with a diffusion model
WHAT IT IS
Stable Video Diffusion (SVD) is an image-to-video model Stability AI released in November 2023 with open weights. It frames image-to-video generation as latent-space diffusion: conditioned on a single still image, it produces a short clip with an adjustable amount of motion. The model generates 14 frames at 576×1024 resolution by default. It shipped as both an image-to-video model and a multi-view model, the latter synthesising views orbiting an object from a single image.
Why it matters
It was one of the few image-to-video models to release open weights, letting “make a short video from one image” run locally or on self-hosted services and advancing open practice in image-to-video.
Key specs
- Resolution
- 576×1024
- Default frames
- 14 frames
- Input
- Single image
- Open weights
- Yes
Capabilities
Related concepts
Latent Diffusion & Conditional Control
Run diffusion not over pixels, but inside a compressed semantic space
Diffusion Models
Learn a thousand tiny denoising steps, and you can build an image from pure noise
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world
Comparable products
Sora
2024Generates coherent video up to about a minute long from a description
Gen-3 Alpha
2024A highly controllable text-to-video model for film and advertising
Kling
2024A short-video model for both text-to-video and image-to-video
Dream Machine
2024Generates short videos with motion from text or an image