Instructions to use HelloSun/SDXL-base-1.0-OpenVINO-INT4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use HelloSun/SDXL-base-1.0-OpenVINO-INT4 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("HelloSun/SDXL-base-1.0-OpenVINO-INT4", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
SDXL base 1.0 OpenVINO INT4
SDXL base 1.0 OpenVINO INT4
OpenVINO + NNCF weight-only INT4 conversion of stabilityai/stable-diffusion-xl-base-1.0.
unet + text_encoder_2 (OpenCLIP bigG):
OVWeightQuantizationConfig(bits=4, sym=False, group_size=128, group_size_fallback="adjust", ratio=1.0)
rest: default INT8 (OVWeightQuantizationConfig() bits=8)
via OVQuantizer + ov_config=OVConfig(quantization_config=OVPipelineQuantizationConfig(...))
FP16 export: optimum-cli export openvino -m stabilityai/stable-diffusion-xl-base-1.0 --task text2image-stable-diffusion --variant fp16 --weight-format fp16 sdxl-ov-fp16
Inference params follow the original SDXL base 1.0 model card: 1024x1024, num_inference_steps=50, guidance_scale=7.0.
(The sibling HelloSun/FLUX.2-klein-4B-OpenVINO-INT4 uses 4 steps / guidance 1.0 because FLUX.2-klein-4B is a 4-step distilled model. Do not copy those values here.)
Pinned env: diffusers==0.37.1 transformers==4.57.6 tokenizers==0.22.0 huggingface-hub==0.35.1 optimum-intel==2.2.0 optimum==2.3.0 openvino==2026.4.0 nncf==3.4.0 torch pillow psutil accelerate
License. This is a derivative of
stabilityai/stable-diffusion-xl-base-1.0, which is released under the CreativeML Open RAIL++-M License (license: other), not Apache-2.0. Anyone redistributing or serving this repo must keep complying with the original Stability AI terms, including the use-based restrictions in Attachment A. See LICENSE.md. The base model is ungated, so no license acceptance is required to download it.
Examples (1024x1024, INT4, CPU)
01_hanfu (seed 42)
Young Chinese woman in red Hanfu, intricate embroidery, impeccable makeup, red floral forehead pattern, elaborate high bun, golden phoenix headdress, soft-lit outdoor night background, silhouetted tiered pagoda, blurred colorful distant lights, photorealistic, ultra detailed, 8k
02_astronaut (seed 43)
Astronaut in a jungle, cold color palette, muted colors, detailed, 8k, photorealistic, cinematic lighting
03_taipei (seed 44)
Cyberpunk street in Taipei at night, heavy rain, neon signs with text 'TAIPEI' and Chinese characters '台北', reflections on wet asphalt, crowded night market, cinematic, ultra detailed
04_shiba (seed 45)
Cute Shiba Inu wearing a tiny astronaut helmet, sitting in a field of sunflowers under a starry sky, dreamy illustration, vibrant colors, high quality
05_ink (seed 46)
Traditional Chinese ink wash landscape, misty mountains, a small pagoda on a cliff, cranes flying, minimalist, elegant, high aesthetic quality
512px resized controls are in outputs/ (*_512.png), full timing in outputs/benchmark.json.
實驗數據
HW: Intel Xeon Platinum 8559C, 192 logical cores, OpenVINO 2026.4.0. Load+compile 12.35s, final RSS 3491.0MB.
| name | seed | steps | guidance | 總時間(s) | 平均單步(s) | peak RSS(MB) |
|---|---|---|---|---|---|---|
| 01_hanfu | 42 | 50 | 7.0 | 188.30 | 3.77 | 3491.0 |
| 02_astronaut | 43 | 50 | 7.0 | 192.58 | 3.85 | 3491.0 |
| 03_taipei | 44 | 50 | 7.0 | 184.58 | 3.69 | 3491.0 |
| 04_shiba | 45 | 50 | 7.0 | 213.98 | 4.28 | 3491.0 |
| 05_ink | 46 | 50 | 7.0 | 238.05 | 4.76 | 3491.0 |
Quantize: 64.68s. Sizes: FP16 unet 4.80GB / text_encoder_2 1.30GB (6.50GB total) -> INT4 unet 1.40GB / text_encoder_2 0.37GB (1.97GB total, ~3.30x).
詳見 REPORT.md + outputs/benchmark.json + outputs/prompts.txt.
用法
pip install diffusers==0.37.1 transformers==4.57.6 tokenizers==0.22.0 huggingface-hub==0.35.1 \
optimum-intel==2.2.0 optimum==2.3.0 openvino==2026.4.0 nncf==3.4.0 torch pillow psutil accelerate
from optimum.intel import OVDiffusionPipeline
import torch
pipe = OVDiffusionPipeline.from_pretrained("HelloSun/SDXL-base-1.0-OpenVINO-INT4", compile=True)
img = pipe(
prompt="Young Chinese woman in red Hanfu, intricate embroidery, ...",
num_inference_steps=50,
guidance_scale=7.0,
height=1024,
width=1024,
generator=torch.Generator().manual_seed(42),
).images[0]
img.save("out.png")
或用本 repo 腳本:
python inference_int4.py --prompt "Astronaut in a jungle, cold color palette, ..." --seed 43 --output out.png
python generate5.py # 複現 5 張 + benchmark.json/prompts.txt
python quantize_int4.py # FP16 導出 + INT4 量化(見 REPORT.md)
Speed note. 50 steps at 1024x1024 is the original model card's setting, and it is slow on CPU. For a faster
preview: --steps 30, or --width/--height 512. Lowering steps below ~30 visibly degrades SDXL base 1.0; it is not a
distilled few-step model like FLUX.2-klein-4B.
檔案結構
./ (INT4 模型: model_index.json, unet/, text_encoder/, text_encoder_2/, vae_decoder/,
vae_encoder/, scheduler/, tokenizer/, tokenizer_2/, openvino_config.json)
README.md / REPORT.md
inference_int4.py / quantize_int4.py / generate5.py
examples/01_hanfu.png ... 05_ink.png (展示用, 同 outputs 1024)
outputs/benchmark.json / prompts.txt / benchmark_quantization.json / *_1024.png / *_512.png
轉換細節
The base model is a StableDiffusionXLPipeline with two text encoders:
| component | class | approx. params | precision used here |
|---|---|---|---|
unet |
UNet2DConditionModel |
2.57B | INT4 |
text_encoder_2 |
CLIPTextModelWithProjection (OpenCLIP bigG) |
694M | INT4 |
text_encoder |
CLIPTextModel (CLIP ViT-L) |
123M | INT8 |
vae_decoder / vae_encoder |
AutoencoderKL |
84M | INT8 |
The UNet carries the vast majority of both compute and memory, so it is the primary INT4 target. Of the two text
encoders, text_encoder_2 (OpenCLIP bigG) holds most of the prompt semantics and is ~5x the size of CLIP ViT-L, so
it is the one worth compressing; the small CLIP-L encoder stays INT8, as do the VAE decoder/encoder, because per the
optimum-intel docs "quantizing the rest of the diffusion pipeline does not significantly improve inference
performance but could potentially lead to substantial accuracy degradation."
group_size_fallback="adjust" matters here: several UNet projection shapes are not divisible by 128, and without it
NNCF would silently skip those layers and leave them in FP16. See REPORT.md.
Model tree for HelloSun/SDXL-base-1.0-OpenVINO-INT4
Base model
stabilityai/stable-diffusion-xl-base-1.0