SDXL base 1.0 OpenVINO INT4

SDXL base 1.0 OpenVINO INT4

OpenVINO + NNCF weight-only INT4 conversion of stabilityai/stable-diffusion-xl-base-1.0.

unet + text_encoder_2 (OpenCLIP bigG): OVWeightQuantizationConfig(bits=4, sym=False, group_size=128, group_size_fallback="adjust", ratio=1.0)

rest: default INT8 (OVWeightQuantizationConfig() bits=8)

via OVQuantizer + ov_config=OVConfig(quantization_config=OVPipelineQuantizationConfig(...))

FP16 export: optimum-cli export openvino -m stabilityai/stable-diffusion-xl-base-1.0 --task text2image-stable-diffusion --variant fp16 --weight-format fp16 sdxl-ov-fp16

Inference params follow the original SDXL base 1.0 model card: 1024x1024, num_inference_steps=50, guidance_scale=7.0. (The sibling HelloSun/FLUX.2-klein-4B-OpenVINO-INT4 uses 4 steps / guidance 1.0 because FLUX.2-klein-4B is a 4-step distilled model. Do not copy those values here.)

Pinned env: diffusers==0.37.1 transformers==4.57.6 tokenizers==0.22.0 huggingface-hub==0.35.1 optimum-intel==2.2.0 optimum==2.3.0 openvino==2026.4.0 nncf==3.4.0 torch pillow psutil accelerate

License. This is a derivative of stabilityai/stable-diffusion-xl-base-1.0, which is released under the CreativeML Open RAIL++-M License (license: other), not Apache-2.0. Anyone redistributing or serving this repo must keep complying with the original Stability AI terms, including the use-based restrictions in Attachment A. See LICENSE.md. The base model is ungated, so no license acceptance is required to download it.


Examples (1024x1024, INT4, CPU)

01_hanfu (seed 42)

01_hanfu Young Chinese woman in red Hanfu, intricate embroidery, impeccable makeup, red floral forehead pattern, elaborate high bun, golden phoenix headdress, soft-lit outdoor night background, silhouetted tiered pagoda, blurred colorful distant lights, photorealistic, ultra detailed, 8k

02_astronaut (seed 43)

02_astronaut Astronaut in a jungle, cold color palette, muted colors, detailed, 8k, photorealistic, cinematic lighting

03_taipei (seed 44)

03_taipei Cyberpunk street in Taipei at night, heavy rain, neon signs with text 'TAIPEI' and Chinese characters '台北', reflections on wet asphalt, crowded night market, cinematic, ultra detailed

04_shiba (seed 45)

04_shiba Cute Shiba Inu wearing a tiny astronaut helmet, sitting in a field of sunflowers under a starry sky, dreamy illustration, vibrant colors, high quality

05_ink (seed 46)

05_ink Traditional Chinese ink wash landscape, misty mountains, a small pagoda on a cliff, cranes flying, minimalist, elegant, high aesthetic quality

512px resized controls are in outputs/ (*_512.png), full timing in outputs/benchmark.json.


實驗數據

HW: Intel Xeon Platinum 8559C, 192 logical cores, OpenVINO 2026.4.0. Load+compile 12.35s, final RSS 3491.0MB.

name seed steps guidance 總時間(s) 平均單步(s) peak RSS(MB)
01_hanfu 42 50 7.0 188.30 3.77 3491.0
02_astronaut 43 50 7.0 192.58 3.85 3491.0
03_taipei 44 50 7.0 184.58 3.69 3491.0
04_shiba 45 50 7.0 213.98 4.28 3491.0
05_ink 46 50 7.0 238.05 4.76 3491.0

Quantize: 64.68s. Sizes: FP16 unet 4.80GB / text_encoder_2 1.30GB (6.50GB total) -> INT4 unet 1.40GB / text_encoder_2 0.37GB (1.97GB total, ~3.30x). 詳見 REPORT.md + outputs/benchmark.json + outputs/prompts.txt.


用法

pip install diffusers==0.37.1 transformers==4.57.6 tokenizers==0.22.0 huggingface-hub==0.35.1 \
            optimum-intel==2.2.0 optimum==2.3.0 openvino==2026.4.0 nncf==3.4.0 torch pillow psutil accelerate
from optimum.intel import OVDiffusionPipeline
import torch

pipe = OVDiffusionPipeline.from_pretrained("HelloSun/SDXL-base-1.0-OpenVINO-INT4", compile=True)
img = pipe(
    prompt="Young Chinese woman in red Hanfu, intricate embroidery, ...",
    num_inference_steps=50,
    guidance_scale=7.0,
    height=1024,
    width=1024,
    generator=torch.Generator().manual_seed(42),
).images[0]
img.save("out.png")

或用本 repo 腳本:

python inference_int4.py --prompt "Astronaut in a jungle, cold color palette, ..." --seed 43 --output out.png
python generate5.py            # 複現 5 張 + benchmark.json/prompts.txt
python quantize_int4.py        # FP16 導出 + INT4 量化(見 REPORT.md)

Speed note. 50 steps at 1024x1024 is the original model card's setting, and it is slow on CPU. For a faster preview: --steps 30, or --width/--height 512. Lowering steps below ~30 visibly degrades SDXL base 1.0; it is not a distilled few-step model like FLUX.2-klein-4B.


檔案結構

./ (INT4 模型: model_index.json, unet/, text_encoder/, text_encoder_2/, vae_decoder/,
     vae_encoder/, scheduler/, tokenizer/, tokenizer_2/, openvino_config.json)
README.md / REPORT.md
inference_int4.py / quantize_int4.py / generate5.py
examples/01_hanfu.png ... 05_ink.png   (展示用, 同 outputs 1024)
outputs/benchmark.json / prompts.txt / benchmark_quantization.json / *_1024.png / *_512.png

轉換細節

The base model is a StableDiffusionXLPipeline with two text encoders:

component class approx. params precision used here
unet UNet2DConditionModel 2.57B INT4
text_encoder_2 CLIPTextModelWithProjection (OpenCLIP bigG) 694M INT4
text_encoder CLIPTextModel (CLIP ViT-L) 123M INT8
vae_decoder / vae_encoder AutoencoderKL 84M INT8

The UNet carries the vast majority of both compute and memory, so it is the primary INT4 target. Of the two text encoders, text_encoder_2 (OpenCLIP bigG) holds most of the prompt semantics and is ~5x the size of CLIP ViT-L, so it is the one worth compressing; the small CLIP-L encoder stays INT8, as do the VAE decoder/encoder, because per the optimum-intel docs "quantizing the rest of the diffusion pipeline does not significantly improve inference performance but could potentially lead to substantial accuracy degradation."

group_size_fallback="adjust" matters here: several UNet projection shapes are not divisible by 128, and without it NNCF would silently skip those layers and leave them in FP16. See REPORT.md.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for HelloSun/SDXL-base-1.0-OpenVINO-INT4

Finetuned
(1220)
this model