本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
これは何か
Stable Diffusion is an open-source text-to-image model that Stability AI released in August 2022. It runs the diffusion process in the latent space of a pretrained autoencoder, sharply cutting computation so the model can run on consumer GPUs. Text is turned into conditioning vectors by a CLIP text encoder and injected into the denoising network through cross-attention. The first 1.x version natively outputs 512×512, with a U-Net of roughly 860 million parameters and weights released under an open licence.
なぜ覚えておく価値があるか
By open-sourcing text-to-image weights and making them runnable on consumer GPUs, it directly spawned the whole open-source image ecosystem of fine-tunes, plugins and local generation, marking the point where generative imagery reached a broad audience.
主な仕様
- Native resolution
- 512×512 (1.x)
- U-Net parameters
- About 860M
- Architecture
- Latent diffusion + CLIP text encoder
- Open weights
- Yes
- Released
- 2022-08
対応する能力
関連する概念
同種の製品
Midjourney
2022美的スタイルで知られるテキストから画像生成サービス
DALL·E 3
2023長い指示を詳細な説明に書き換えてから画像を生成する
FLUX
2024整流フローTransformerで高解像度画像を生成する
Imagen
2022カスケード拡散と大規模テキストエンコーダで写実的な画像を生成する
Firefly
2023クリエイター向けの画像生成・編集ツール
Seedream
2024ネイティブ高解像度で文字描画に強いテキスト画像生成モデル