Stable Diffusion

Stable Diffusion is a family of text-conditioned image-generation models built around latent diffusion. Instead of denoising full-resolution pixels directly, the model compresses an image into a lower-dimensional latent representation, denoises that latent under text conditioning, and decodes the final latent back to pixels. That latent-space design is the practical reason Stable Diffusion-style systems can produce high-resolution images with less compute than pixel-space diffusion.

Stable Diffusion sits between generative AI and computer vision. The generator is a diffusion model, but it depends on visual representation learning: an autoencoder defines the image latent space, a text or vision-language encoder supplies conditioning, and the denoiser learns visual structure from large image corpora. For representation-side context, see self-supervised visual learning.

Latent diffusion

A latent diffusion model first encodes an image into a latent:

where is usually an autoencoder encoder. The forward diffusion process adds Gaussian noise:

The denoising model receives the noisy latent , timestep , and conditioning from a prompt encoder, then predicts the noise:

At sampling time, the model starts from noise and repeatedly applies a scheduler step using until it obtains a clean latent . A decoder maps that latent back to an image:

Stable Diffusion denoises a compressed latent under text conditioning, then decodes the final latent back to pixels.

Guidance

Text conditioning is commonly strengthened with classifier-free guidance. The denoiser is trained sometimes with the text condition and sometimes without it. During sampling, the two predictions are combined:

Here is the guidance scale. Larger usually makes the image follow the prompt more strongly, but it can reduce diversity, over-sharpen textures, or amplify artifacts. A negative prompt is an engineering variant: replace the empty condition with a condition describing what the sample should move away from.

Architecture Variants

componentclassic latent-diffusion Stable Diffusionlater variants
Image spaceAutoencoder compresses pixels into a spatial latent and decodes final latents back to pixels.The latent-space principle remains common, though autoencoder details change.
DenoiserU-Net with residual blocks, attention, timestep embedding, and cross-attention to text.Larger U-Nets in SDXL; diffusion-transformer or rectified-flow backbones in newer systems.
ConditioningText encoder produces prompt embeddings exposed through cross-attention.Multiple text encoders, richer conditioning, image conditioning, control maps, or multimodal token mixing.
SamplerIterative denoising schedule over timesteps.Faster samplers, distillation, consistency-style methods, or rectified-flow trajectories.

The important conceptual point is stable across variants: generation is an iterative denoising process in a learned visual latent space, steered by a conditioning signal.

Worked sampling scenario

For a prompt such as “a watercolor sketch of a glass greenhouse at sunrise,” the system follows this path:

stagerepresentationrole
Prompt encodingText embeddings Encodes concepts such as watercolor, greenhouse, glass, and sunrise.
Initial latentRandom noise Provides stochastic variation; different seeds start from different noise.
Denoising loopLatents Repeatedly removes noise while cross-attention steers the latent toward the prompt.
GuidanceTrades prompt adherence against diversity and artifact risk.
DecodingImage Converts the final latent into RGB pixels.

Image-to-image and inpainting use the same mechanism with a different starting point: encode an existing image, add a controlled amount of noise, and denoise under the prompt or mask constraints. More noise gives the model more freedom; less noise preserves more of the original image.

Caveats

Stable Diffusion is not a factual image database. It can invent details, reproduce dataset biases, struggle with exact text, count objects poorly, or produce anatomically inconsistent results. Prompt adherence, aesthetic quality, and diversity trade off through sampling settings and guidance. For product use, evaluate copyright/licensing constraints, safety filters, demographic bias, prompt injection through image-editing workflows, and whether generated images need provenance or watermarking.

References