Instructions to use zeromodels/stable-diffusion-2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- ZeroModels
How to use zeromodels/stable-diffusion-2 with ZeroModels:
# pip install -U zeromodels # ZeroModels is pure Keras 3, so pick a backend: "jax", "torch" or "tensorflow". import os os.environ["KERAS_BACKEND"] = "jax" from zeromodels import AutoZModel # AutoZModel reads the repo's model_type and loads the matching class. # For a task head use the matching loader, e.g. AutoZMImageClassify / AutoZMDetect / # AutoZMSemanticSegment / AutoZMTextGenerate (see zeromodels.auto). model = AutoZModel.from_weights("zeromodels/stable-diffusion-2") - Keras
How to use zeromodels/stable-diffusion-2 with Keras:
# Available backend options are: "jax", "torch", "tensorflow". import os os.environ["KERAS_BACKEND"] = "jax" import keras model = keras.saving.load_model("hf://zeromodels/stable-diffusion-2") - Notebooks
- Google Colab
- Kaggle
See our collection for all Stable Diffusion 2.x checkpoints.
Run Stable Diffusion 2.x with Keras 3: JAX, PyTorch, or TensorFlow
zeromodels/stable-diffusion-2
Paper: High-Resolution Image Synthesis with Latent Diffusion Models (arXiv:2112.10752) | HF Papers
Pure-Keras 3 conversion of sd2-community/stable-diffusion-2 for
zeromodels. One implementation runs unmodified on
TensorFlow / Torch / JAX. The whole text-to-image model ships as one container:
the UNet denoiser, the VAE and the OpenCLIP ViT-H/14 text encoder (penultimate layer) in model.weights.h5
(1.29B parameters, 4.81 GB), plus zm_config.json (the three
component configs, the checkpoint's DDIMScheduler schedule with its v_prediction objective
and the default generation settings) and the tokenizer as tokenizer.json. Weights are stored in float32, exactly as released.
This checkpoint generates 768x768 images (a 96x96 latent).
For model details, intended use and limitations, see the upstream model card.
Architecture
| Component | zeromodels class | Details |
|---|---|---|
| Denoiser | UNet2DConditionModel |
(320, 640, 1280, 1280) channels, 2 ResNet blocks per level, (5, 10, 20, 20) attention heads on the 1024-d text context, linear token projection, 96x96x4 latent |
| Autoencoder | AutoencoderKL |
(128, 256, 512, 512) channels, x8 spatial compression to 4 latent channels, scaling_factor 0.18215 |
| Text encoder | CLIPTextModel |
OpenCLIP ViT-H/14 text encoder (penultimate layer): 1024-d, 23 layers, 16 heads, 77 tokens, gelu |
| Scheduler | DDIMScheduler |
scaled_linear betas 0.00085 to 0.012 over 1000 steps, v_prediction; DDIM / PNDM / Euler / Euler-ancestral are drop-in |
Quick start
import os
os.environ["KERAS_BACKEND"] = "torch" # or "jax" / "tensorflow"
from PIL import Image
from zeromodels.models.stable_diffusion_2 import StableDiffusion2TextToImage, StableDiffusion2Tokenizer
model = StableDiffusion2TextToImage.from_weights("zeromodels/stable-diffusion-2")
tokenizer = StableDiffusion2Tokenizer.from_weights("zeromodels/stable-diffusion-2")
inputs = tokenizer("a photograph of an astronaut riding a horse")
images = model.generate(**inputs, num_inference_steps=50, guidance_scale=7.5, seed=0)
Image.fromarray(images[0]).save("astronaut.png") # (768, 768, 3) uint8
generate takes the tokenizer's input_ids (batch them for several prompts), an optional
negative_input_ids (tokenize the negative prompt), num_inference_steps, guidance_scale,
a seed, or explicit latents of shape (batch, 96, 96, 4) for results that are
identical across backends.
Load any Stable Diffusion 2.x checkpoint the same way with from_weights("zeromodels/<variant>"):
| Variant | Hub | Training |
|---|---|---|
stable-diffusion-2-base |
zeromodels/stable-diffusion-2-base | 512px, epsilon: from scratch, 550k steps at 256px on LAION-5B (aesthetics >= 4.5), then 850k steps at 512px |
stable-diffusion-2 |
zeromodels/stable-diffusion-2 | 768px, v-prediction: 2-base + 150k steps at 768px |
stable-diffusion-2-1-base |
zeromodels/stable-diffusion-2-1-base | 512px, epsilon: 2-base + 220k steps at 512px (punsafe 0.98) |
stable-diffusion-2-1 |
zeromodels/stable-diffusion-2-1 | 768px, v-prediction: 2 + 55k steps (punsafe 0.1) + 155k steps (punsafe 0.98) at 768px |
sd-turbo |
zeromodels/sd-turbo | 512px, epsilon, Euler (trailing spacing), 1 to 4 steps, no guidance: SD 2.1 distilled with Adversarial Diffusion Distillation (Stability AI Community License) |
Tips
Set
KERAS_BACKENDbefore importing Keras / zeromodels.The graphs are built for 768px. Pass
unet_sample_size=<px / 8>, vae_sample_size=<px>tofrom_weightsto build for another multiple of 64px (the weights are resolution-independent).Swap the sampler any time:
model.scheduler = EulerDiscreteScheduler.from_config(model.config.scheduler_config)(zeromodels.base.base_scheduler).StableDiffusion2Model.from_weights(...)loads the same repo as the bare container (UNet / VAE / text encoder as.unet/.vae/.text_encoder) without the generation loop.Both
channels_lastandchannels_firstare supported (keras.config.set_image_data_formatbefore loading);generatealways returns(batch, H, W, 3)uint8.On-the-fly
hf:conversion is not supported for diffusion models; the checkpoints are hosted here, converted once.See the Stable Diffusion 2.x docs.
License
The weights are redistributed under the CreativeML Open RAIL++-M License of the upstream checkpoint, including its use-based restrictions. By using them you agree to those terms.
Notice
Modifications by zeromodels (https://github.com/IMvision12/ZeroModels): the checkpoint
released at https://huggingface.co/sd2-community/stable-diffusion-2 was converted to the
Keras 3 weights layout of zeromodels (model.weights.h5, zm_config.json, tokenizer.json),
stored in float32 as released. The model architecture and the parameter values are
unchanged; the weight names and the file format differ from the release.
Special Thanks
Thank you to Stability AI and the LAION / OpenCLIP teams for training and releasing Stable Diffusion, and to the Hugging Face diffusers team, whose implementation this port was verified against.
- Downloads last month
- -
Model tree for zeromodels/stable-diffusion-2
Base model
sd2-community/stable-diffusion-2