이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
무엇인가
Stable Diffusion is an open-source text-to-image model that Stability AI released in August 2022. It runs the diffusion process in the latent space of a pretrained autoencoder, sharply cutting computation so the model can run on consumer GPUs. Text is turned into conditioning vectors by a CLIP text encoder and injected into the denoising network through cross-attention. The first 1.x version natively outputs 512×512, with a U-Net of roughly 860 million parameters and weights released under an open licence.
기억할 만한 이유
By open-sourcing text-to-image weights and making them runnable on consumer GPUs, it directly spawned the whole open-source image ecosystem of fine-tunes, plugins and local generation, marking the point where generative imagery reached a broad audience.
주요 사양
- Native resolution
- 512×512 (1.x)
- U-Net parameters
- About 860M
- Architecture
- Latent diffusion + CLIP text encoder
- Open weights
- Yes
- Released
- 2022-08
소속 능력
관련 개념
동종 제품
Midjourney
2022미적 스타일로 알려진 텍스트-이미지 서비스
DALL·E 3
2023긴 프롬프트를 상세한 설명으로 다시 쓴 뒤 이미지를 생성한다
FLUX
2024정류 흐름 트랜스포머로 고해상도 이미지를 생성한다
Imagen
2022계단식 확산과 대형 텍스트 인코더로 사실적인 이미지를 생성한다
Firefly
2023크리에이터를 위한 이미지 생성·편집 도구
Seedream
2024네이티브 고해상도와 뛰어난 문자 렌더링의 텍스트-이미지 모델