Stable Diffusion
Released text-to-image weights openly and small enough to run on consumer GPUs
WHAT IT IS
Stable Diffusion is an open-source text-to-image model that Stability AI released in August 2022. It runs the diffusion process in the latent space of a pretrained autoencoder, sharply cutting computation so the model can run on consumer GPUs. Text is turned into conditioning vectors by a CLIP text encoder and injected into the denoising network through cross-attention. The first 1.x version natively outputs 512×512, with a U-Net of roughly 860 million parameters and weights released under an open licence.
Why it matters
By open-sourcing text-to-image weights and making them runnable on consumer GPUs, it directly spawned the whole open-source image ecosystem of fine-tunes, plugins and local generation, marking the point where generative imagery reached a broad audience.
Key specs
- Native resolution
- 512×512 (1.x)
- U-Net parameters
- About 860M
- Architecture
- Latent diffusion + CLIP text encoder
- Open weights
- Yes
- Released
- 2022-08
Capabilities
Related concepts
Latent Diffusion & Conditional Control
Run diffusion not over pixels, but inside a compressed semantic space
Diffusion Models
Learn a thousand tiny denoising steps, and you can build an image from pure noise
Autoencoders & VAE
Squeeze information through a bottleneck, then let it grow back
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world
Comparable products
Midjourney
2022A text-to-image service known for its aesthetic style
DALL·E 3
2023Rewrites a long prompt into a detailed description, then draws the image
FLUX
2024Generates high-resolution images with a rectified-flow transformer
Imagen
2022Generates photorealistic images with cascaded diffusion and a large text encoder
Firefly
2023An image generation and editing tool aimed at creators
Seedream
2024A text-to-image model with native high resolution and strong text rendering