Juggernaut-XL-Lightning · OpenVINO INT4 (CPU)
RunDiffusion/Juggernaut-XL-Lightning 轉成 OpenVINO IR,
unet(transformer) 與 text_encoder 以 NNCF 做 weight-only INT4(asymmetric, group_size 128),
其餘組件 INT8。不需要 GPU,純 CPU 即可生圖。
- 權重體積 6.65 GB → 2.27 GB(-65.8%)
- 1024×1024 / 6 steps / CFG 1.8:**~23.6 s 出一張(3.6 s/step)**(Xeon Platinum 8559C, 16 vCPU)
- 512×512 / 6 steps:**~7.5 s 出一張(1.17 s/step)**
- load + compile:12.8 s
轉換環境:openvino 2026.4.0 / nncf 3.4.0 / optimum 2.3.0 / optimum-intel 2.2.0 / diffusers 0.39.0 / transformers 5.5.4 / tokenizers 0.22.2 / huggingface-hub 1.21.0 / torch 2.14.1 / pillow 12.3.0 / psutil 7.2.2
模型頁展示(10 張)
1024 × 1024(6 steps, DPM++ 2M Karras, CFG 1.8)
512 × 512 對照組(同 prompt / 同 seed)
Prompts 見 outputs/prompts.txt。
安裝
pip install -U optimum optimum-intel openvino nncf \
diffusers transformers tokenizers huggingface-hub \
torch pillow psutil
用法
import torch
from optimum.intel import OVDiffusionPipeline
pipe = OVDiffusionPipeline.from_pretrained(
"HelloSun/Juggernaut-XL-Lightning-OpenVINO-INT4",
compile=True,
device="CPU",
)
prompt = (
"Young Chinese woman in red Hanfu, intricate embroidery, impeccable makeup, "
"red floral forehead pattern, elaborate high bun, golden phoenix headdress, "
"soft-lit outdoor night background, silhouetted tiered pagoda, blurred colorful "
"distant lights, photorealistic, ultra detailed, 8k"
)
generator = torch.Generator(device="cpu").manual_seed(42)
image = pipe(
prompt=prompt,
negative_prompt="lowres, worst quality, low quality, blurry, jpeg artifacts, watermark",
width=1024, height=1024,
num_inference_steps=6,
guidance_scale=1.8,
generator=generator,
).images[0]
image.save("hanfu.png")
完整腳本:
| 檔案 | 用途 |
|---|---|
inference_int4.py |
單張 txt2img / img2img 推論 |
quantize_int4.py |
由 FP16 pipeline 產生 INT4 pipeline(OVQuantizer + NNCF) |
generate5_xl_lightning.py |
5 組固定 seed 批次生成,callback 記錄每 step 時間與 RSS |
REPORT.md |
完整轉換與實驗報告 |
outputs/benchmark.json |
機器資訊 + 每 step 時間/RSS 原始數據 |
outputs/prompts.txt |
全部 prompt / negative prompt / seed |
轉換步驟(重現本 repo)
# 1) 匯出 FP16
optimum-cli export openvino -m RunDiffusion/Juggernaut-XL-Lightning \
--task text-to-image --library diffusers --weight-format fp16 \
./juggernaut-xl-lightning-ov-fp16
# 2) NNCF weight-only INT4(unet + text_encoder INT4,其餘 INT8)
python quantize_int4.py \
--model_path ./juggernaut-xl-lightning-ov-fp16 \
--output_path ./juggernaut-xl-lightning-ov-int4
量化設定(quantize_int4.py 內):
int4 = dict(bits=4, sym=False, group_size=128, group_size_fallback="adjust", ratio=1.0)
OVPipelineQuantizationConfig(
quantization_configs={
"unet": OVWeightQuantizationConfig(**int4),
"text_encoder": OVWeightQuantizationConfig(**int4),
},
default_config=OVWeightQuantizationConfig(bits=8),
)
實驗數據
環境:Intel Xeon Platinum 8559C(2 socket / 96 core / 192 thread,容器 cgroup 限 16 vCPU)、
AVX-512、2 TB RAM、無 GPU。lscpu CPU(s)=192,os.cpu_count()=192,OpenVINO 2026.4.0。
| 項目 | 數值 |
|---|---|
load + compile(compile=True) |
12.80 s(結束時 RSS 3399 MB) |
| 1024×1024 / 6 steps 單張總時間 | 22.69 – 25.82 s(平均 23.62 s) |
| 1024×1024 單步時間 | 平均 3.62 s、中位數 3.46 s、範圍 3.29 – 4.15 s(暖機除外) |
| 512×512 / 6 steps 單張總時間 | 7.32 – 7.67 s(平均 7.50 s) |
| 512×512 單步時間 | 平均 1.17 s、中位數 1.11 s |
| 峰值 RSS(連續生成) | ~10.8 GB |
| FP16 pipeline 大小 | 6635.4 MB |
| INT4 pipeline 大小 | 2268.7 MB(-65.8%) |
| 量化時間(weight-only, data-free) | 50.3 s |
逐張明細:
| 影像 | seed | 解析度 | steps | CFG | 總時間 | 單步平均 | 峰值 RSS |
|---|---|---|---|---|---|---|---|
| 01_hanfu | 42 | 1024² | 6 | 1.8 | 25.82 s | 3.919 s | 8510 MB |
| 02_astronaut | 43 | 1024² | 6 | 1.8 | 23.53 s | 3.616 s | 10781 MB |
| 03_taipei | 44 | 1024² | 6 | 1.8 | 22.69 s | 3.479 s | 10784 MB |
| 04_shiba | 45 | 1024² | 6 | 1.8 | 22.95 s | 3.527 s | 10786 MB |
| 05_ink | 46 | 1024² | 6 | 1.8 | 23.11 s | 3.546 s | 10788 MB |
| 01_hanfu_512 | 42 | 512² | 6 | 1.8 | 7.67 s | 1.181 s | 10826 MB |
| 02_astronaut_512 | 43 | 512² | 6 | 1.8 | 7.61 s | 1.189 s | 10828 MB |
| 03_taipei_512 | 44 | 512² | 6 | 1.8 | 7.32 s | 1.151 s | 10830 MB |
| 04_shiba_512 | 45 | 512² | 6 | 1.8 | 7.48 s | 1.150 s | 10832 MB |
| 05_ink_512 | 46 | 512² | 6 | 1.8 | 7.41 s | 1.154 s | 10834 MB |
建議參數
沿用上游模型卡建議:sampler DPM++ SDE Karras、steps 5–7、CFG 1.5–2.0(越低越寫實)。
本 repo 的 scheduler_config.json 已設為 DPMSolverMultistepScheduler + use_karras_sigmas=true(對應 2M Karras,在 Lightning 上效果接近 SDE Karras),
model_index.json["scheduler"] 也同步指向它,開箱即對應。
Credits
- 模型:RunDiffusion/Juggernaut-XL-Lightning by RunDiffusion / KandooAI(CreativeML Open RAIL-M)
- 轉換與範例:HelloSun/Juggernaut-XL-Lightning-OpenVINO-INT4
Model tree for HelloSun/Juggernaut-XL-Lightning-OpenVINO-INT4
Base model
stabilityai/stable-diffusion-xl-base-1.0








