Stable Diffusion 3.5 Medium — OpenVINO INT4
將 stabilityai/stable-diffusion-3.5-medium 轉換為 OpenVINO INT4 權重,並以 純 CPU 實測出圖。
optimum-intel 匯出 → NNCF weight-only INT4 量化(耗時 69 s)→ FP16 15.17 GB → INT4 5.08 GB(↓ ~67%)。
⚠️ 基座模型為 gated:需先在
stabilityai/stable-diffusion-3.5-medium頁面接受授權,才能執行轉換腳本下載權重。
目錄
快速資訊
| 項目 | 內容 |
|---|---|
| 基座模型 | stabilityai/stable-diffusion-3.5-medium |
| Pipeline | StableDiffusion3Pipeline |
| Scheduler | FlowMatchEulerDiscreteScheduler(shift=3.0) |
| 量化格式 | transformer、text_encoder_3 → INT4;其餘 → INT8 |
| 模型大小 | FP16 15.17 GB → INT4 5.08 GB(↓ ~67%) |
| 推論裝置 | CPU(OpenVINO CPU plugin,無需 GPU) |
| 推薦參數 | num_inference_steps=40、guidance_scale=4.5、1024×1024 |
| 轉換工具 | optimum-intel 2.2.x / optimum 2.3.0 / OpenVINO 2026.4.0 / NNCF 3.4.0 |
特色
- 三個 text encoder:兩個 CLIP(L / G)維持 INT8,T5-XXL 量化為 INT4。
group_size_fallback="adjust"的關鍵理由:SD3 的 MMDiT projection 維度並非都能被 128 整除,NNCF 預設的"ignore"會靜默跳過這些層、讓它們留在 FP16;改用"adjust"才不會漏量化。- **體積 ↓ 67%**:FP16 15.17 GB → INT4 5.08 GB。
- 記憶體需求低:推論 peak RSS 僅 3.28 GB,是這批 repo 中最低的。
- 純 CPU 可用:不依賴 GPU / CUDA。
- ⚠️ 速度是短板:1024×1024 / 40 steps 平均 226 s / 張。
安裝
pip install diffusers==0.37.1 transformers==4.57.6 tokenizers==0.22.0 huggingface-hub==0.35.1 optimum==2.3.0 optimum-intel==2.2.0 openvino==2026.4.0 nncf==3.4.0 torch pillow psutil sentencepiece accelerate
nncf 僅重新量化時需要。
快速開始
import torch
from optimum.intel import OVDiffusionPipeline
pipe = OVDiffusionPipeline.from_pretrained("HelloSun/SD3.5-medium-OpenVINO-INT4", compile=True)
image = pipe(
prompt="Astronaut in a jungle, cold color palette, muted colors, "
"detailed, 8k, photorealistic, cinematic lighting",
height=1024, width=1024,
num_inference_steps=40,
guidance_scale=4.5,
generator=torch.Generator().manual_seed(43),
).images[0]
image.save("out.png")
完整可執行範例:inference_int4.py 批次生成 + benchmark:generate5.py
推論參數建議
| 參數 | 建議值 | 說明 |
|---|---|---|
num_inference_steps |
40 |
蒸餾 / 推薦步數。請勿隨意增加。 |
guidance_scale |
4.5 |
CFG 設定;Turbo / 蒸餾模型通常為 0.0 或 1.0(即不啟用)。 |
shift |
取自 scheduler/scheduler_config.json |
不需手動傳入,載入時自動套用。 |
height / width |
1024 |
實測解析度。 |
compile=True |
開啟 | 編譯模型以取得較佳效能。 |
範例結果
全部為 40 steps / guidance_scale 4.5 / 1024×1024 / CPU,seed 42–46 固定,可完全重現。
另附 512px 對照圖(
outputs/*_512.png)。
⚠️ 512px 為縮圖:這批 512px 圖是 1024px 輸出的 LANCZOS 縮圖,不是重新以 512px 推理的結果,因此沒有對應的獨立耗時資料。
01_hanfu — seed 42 — 228.4 s
Young Chinese woman in red Hanfu, intricate embroidery, impeccable makeup, red floral forehead pattern, elaborate high bun, golden phoenix headdress, soft-lit outdoor night background, silhouetted tiered pagoda, blurred colorful distant lights, photorealistic, ultra detailed, 8k
02_astronaut — seed 43 — 228.6 s
Astronaut in a jungle, cold color palette, muted colors, detailed, 8k, photorealistic, cinematic lighting
03_taipei — seed 44 — 224.7 s
Cyberpunk street in Taipei at night, heavy rain, neon signs with text 'TAIPEI' and Chinese characters '台北', reflections on wet asphalt, crowded night market, cinematic, ultra detailed
04_shiba — seed 45 — 222.2 s
Cute Shiba Inu wearing a tiny astronaut helmet, sitting in a field of sunflowers under a starry sky, dreamy illustration, vibrant colors, high quality
05_ink — seed 46 — 226.3 s
Traditional Chinese ink wash landscape, misty mountains, a small pagoda on a cliff, cranes flying, minimalist, elegant, high aesthetic quality
風格展示
以下 10 張為 10 類風格 × 同一組 5 組基礎 prompt(seed 42–51), 使用與上方基準測試完全相同的推論設定。
注意:這批圖片沒有對應的
outputs/benchmark.json紀錄,也沒有腳本可重現, 完整 prompt 亦未收錄於 repo(早期說明檔中的版本已被截斷,無法還原)。
traditional_painting · 傳統繪畫
油畫、版畫、水彩等傳統媒材
anime_manga · 動漫繪師
吉卜力 × 新海誠動畫風
digital_3d · 數位 3D
UE5 寫實渲染、CG 質感
photography_cinema · 攝影電影
電影感人像攝影
cultural_regional · 文化地域
敦煌壁畫、水墨等東方美學
material_craft · 材質工藝
絲線刺繡、陶瓷、木雕等
design_commercial · 平面商業
商業插畫、扁平幾何設計
abstract_generative · 抽象生成
分形、生成藝術
era_subculture · 復古次文化
80s 霓虹、賽博龐克
render_technique · 厚塗渲染
油畫厚塗、概念設定
效能實測摘要
完整逐 step 數據見 REPORT.md 與 outputs/benchmark.json。
測試環境
| 項目 | 內容 |
|---|---|
| CPU | Intel(R) Xeon(R) Platinum 8559C |
| 拓撲 | 2 sockets × 48 cores × 2 threads/core = 192 vCPU(96 實體核心) |
| RAM | 2.0 TiB |
| 虛擬化 | KVM(完整虛擬化) |
| OpenVINO | CPU only,2026.4.0(build 2026.4.0-22959-99c81491cc3-releases/2026/4) |
| 設定 | num_inference_steps=40、guidance_scale=4.5、1024×1024 |
總結
| 指標 | 數值 |
|---|---|
| 解析度 | 1024×1024 |
| 平均總耗時 | 226.03 s / 張 |
| 平均單步耗時 | 5.65 s |
| 最快 / 最慢 | 222.18 s / 228.58 s |
| 總計(5 張) | 1130.14 s |
| 記憶體高水位 | 3,276 MB |
- 512px 沒有獨立耗時資料:
outputs/*_512.png是 1024px 輸出的縮圖,非重新推理。 - 文字編碼與 VAE decode 的時間已包含在總耗時內。
逐張結果(1024×1024)
| # | Prompt | Seed | 總耗時 (s) | 平均單步 (s) | guidance_scale |
|---|---|---|---|---|---|
| 01_hanfu | 42 | 228.41 | 5.710 | None | |
| 02_astronaut | 43 | 228.58 | 5.710 | None | |
| 03_taipei | 44 | 224.70 | 5.620 | None | |
| 04_shiba | 45 | 222.18 | 5.550 | None | |
| 05_ink | 46 | 226.27 | 5.660 | None | |
| 平均 | 226.03 | 5.650 |
模型大小
以下為 repo 內 openvino_model.bin 的實際位元組數(Git LFS 記錄值)。
| 元件 | 位元組 | 大小 | 精度 |
|---|---|---|---|
text_encoder_3 |
2,557,956,630 | 2.56 GB | INT4 |
transformer |
1,620,333,792 | 1.62 GB | INT4 |
text_encoder_2 |
696,453,879 | 0.70 GB | — |
text_encoder |
124,172,025 | 0.12 GB | — |
vae_decoder |
49,627,579 | 0.05 GB | — |
vae_encoder |
34,337,212 | 0.03 GB | — |
| 合計 | 5,082,881,117 | 5.08 GB |
FP16 匯出模型約 15.17 GB(transformer 4.60 GB + text_encoder_3 8.87 GB + text_encoder 0.23 GB + text_encoder_2 1.30 GB + vae 0.15 GB)——取自原始轉換紀錄。
INT8 部分(兩個 CLIP text encoder 與 VAE)為 1.72 GB,佔 INT4 總量的 34%。
⚠️ outputs/benchmark_quantization.json 記錄的 int4_size_gb 為 4.76,與 repo 內實際位元組總和 5.08 GB 不符;引用時請以本表為準。
檔案結構
.
├── README.md # 本文件
├── REPORT.md # 完整轉換 + 實測報告
├── requirements.txt # Python 相依套件版本
├── model_index.json # diffusers pipeline 索引(StableDiffusion3Pipeline)
├── openvino_config.json # OpenVINO 量化設定
├── inference_int4.py # 單張推論範例
├── generate5.py # 5 組 prompt 批次生成 + benchmark
├── quantize_int4.py # FP16 匯出 + INT4 量化(可分兩階段執行)
├── transformer/ # INT4 SD3Transformer2DModel(MMDiT 24 層)
├── text_encoder/ # INT8 CLIP-L
├── text_encoder_2/ # INT8 CLIP-G
├── text_encoder_3/ # INT4 T5-XXL
├── tokenizer/ # CLIP tokenizer + OpenVINO tokenizer IR
├── tokenizer_2/ # CLIP tokenizer 2 + OpenVINO tokenizer IR
├── tokenizer_3/ # T5 tokenizer (sentencepiece)
├── vae_encoder/ # INT8 VAE encoder
├── vae_decoder/ # INT8 VAE decoder
├── scheduler/ # FlowMatchEulerDiscreteScheduler (shift=3.0)
├── examples/ # 15 張展示圖(含 10 張風格圖)
└── outputs/ # 10 張實測圖 + 3 個資料檔
├── *_1024.png # 主測組(5 張)
├── *_512.png # 對照組(5 張,為 1024 縮圖)
├── benchmark.json # 逐張耗時 + 硬體/軟體資訊
├── benchmark_quantization.json # 量化設定與耗時
└── prompts.txt # 5 組 prompt 與 seed
從零復現
# 1. 匯出 FP16 OpenVINO 模型
optimum-cli export openvino \
-m stabilityai/stable-diffusion-3.5-medium \
--task text-to-image \
--library diffusers \
--weight-format fp16 \
./SD3.5-medium-ov-fp16
# 2. NNCF weight-only INT4 量化
python quantize_int4.py --fp16-dir ./SD3.5-medium-ov-fp16 \
--int4-dir ./SD3.5-medium-ov-int4
# 3. 單張推論
python inference_int4.py
# 4. 批次生成 5 組 + benchmark
python generate5.py --outdir outputs
量化設定:
from optimum.intel.openvino.configuration import (
OVConfig, OVWeightQuantizationConfig, OVPipelineQuantizationConfig,
)
int4_config = OVWeightQuantizationConfig(
bits=4, sym=False, group_size=128,
group_size_fallback="adjust", ratio=1.0,
)
pipeline_config = OVPipelineQuantizationConfig(
quantization_configs={
"transformer": int4_config,
"text_encoder_3": int4_config,
},
default_config=OVWeightQuantizationConfig(bits=8),
)
已知限制
- 40 steps,CPU 速度慢:1024×1024 平均 226 s / 張(約 3.8 分鐘),是這批 repo 中最慢的。若要快速預覽可降到 28 steps 或改用 512px。
- 512px 對照為縮圖:
outputs/*_512.png是 1024px 輸出的 LANCZOS 縮圖,並非以 512px 重新推理。 - 基座模型為 gated:轉換前必須先接受 Stability AI 的授權條款。
- 授權不是 Apache-2.0:Stability AI Community License 有商用限制,轉載前務必確認條款。
inference_int4.py的--stepsize參數是無效的:該參數定義了但未傳入 pipeline,不會影響任何行為;shift實際上取自scheduler/scheduler_config.json的3.0。outputs/benchmark_quantization.json的int4_size_gb(4.76)與實際位元組總和(5.08 GB)不符。REPORT.md提及的.gitignore不存在於本 repo。
授權與出處
授權:Stability AI Community License(
license: other),不是 Apache-2.0。任何轉載或服務本 repo 都必須繼續遵守 Stability AI 的原始條款。完整條款:LICENSE.md
基座模型為 gated:需先在 Hub 頁面接受授權。
Made with OpenVINO + optimum-intel + NNCF
Model tree for HelloSun/SD3.5-medium-OpenVINO-INT4
Base model
stabilityai/stable-diffusion-3.5-medium













