본문으로 건너뛰기
AI 도감

Stable Diffusion

텍스트-이미지 가중치를 공개하고 소비자용 GPU에서 돌아가게 만들었다

Stability AI 모델 공개 가중치
입력텍스트이미지이미지

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

무엇인가

Stable Diffusion is an open-source text-to-image model that Stability AI released in August 2022. It runs the diffusion process in the latent space of a pretrained autoencoder, sharply cutting computation so the model can run on consumer GPUs. Text is turned into conditioning vectors by a CLIP text encoder and injected into the denoising network through cross-attention. The first 1.x version natively outputs 512×512, with a U-Net of roughly 860 million parameters and weights released under an open licence.

기억할 만한 이유

By open-sourcing text-to-image weights and making them runnable on consumer GPUs, it directly spawned the whole open-source image ecosystem of fine-tunes, plugins and local generation, marking the point where generative imagery reached a broad audience.

주요 사양

Native resolution
512×512 (1.x)
U-Net parameters
About 860M
Architecture
Latent diffusion + CLIP text encoder
Open weights
Yes
Released
2022-08

소속 능력

관련 개념

동종 제품