Instructions to use AcademiaSD/Z-Image-NF4-for-LoRA-Training with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use AcademiaSD/Z-Image-NF4-for-LoRA-Training with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("AcademiaSD/Z-Image-NF4-for-LoRA-Training", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Z-Image NF4 for LoRA Training
Z-Image (the undistilled base model by Tongyi-MAI) quantized to 4-bit NF4 for the Z-Image trainer of AcademiaSD LoRAlab Trainer Studio: LoRA training on consumer NVIDIA GPUs from 8 GB of VRAM.
You do not need to download this repository by hand: the trainer downloads it on first use.
Contents
| Folder | Content | Size |
|---|---|---|
transformer/ |
6B transformer: attention and MLP layers in NF4 (bitsandbytes, pre-quantized), embedders, final layer and adaLN modulation in BF16 | 3.4 GB |
text_encoder/ |
Qwen3-4B text encoder in NF4 (Qwen3Model, without lm_head) |
2.7 GB |
vae/, tokenizer/, scheduler/ |
Unchanged from the original model | 0.2 GB |
| Total | 5.9 GB (original: 20.5 GB) |
The text encoder, tokenizer and VAE are byte-identical in Z-Image and Z-Image-Turbo.
Format
text_encoder/loads withtransformers:Qwen3Model.from_pretrained(".../text_encoder").transformer/is not a standard diffusers checkpoint. It is a pre-quantized cache: one.safetensorsfile per layer inweights/(NF4 weight plus its packed bitsandbytesQuantState, or BF16 for the excluded layers),others.safetensorswith the norms and pad tokens, andindex.json. The trainer builds an emptyZImageTransformer2DModelfromconfig.jsonand fills it from this cache, so the BF16 transformer is never downloaded or loaded.metadata.jsonlists the layers kept in BF16.- The converter that produced it is
tools/zimage/5_conversor_ZImage_NF4.pyin the Trainer Studio repository.
NF4 is meant for training. Images generated with this NF4 transformer are visually equivalent to BF16 (same subject, composition and legible text). For inference, use the original model in ComfyUI.
LoRAs trained with it
The trainer exports LoRAs in diffusers format (transformer.<layer>.lora_A / lora_B plus alpha), which ComfyUI loads directly. Z-Image-Turbo has the same layers, so the LoRAs also load there.
Measured on an RTX 5080 at 512×512, rank 8: ~5.5 GB of VRAM and ~0.9 s per step.
License
Apache 2.0, the same as the original Tongyi-MAI/Z-Image. All credit for the model goes to the Tongyi-MAI team; this repository only changes the storage precision.
Links
- Trainer: AcademiaSD LoRAlab Trainer Studio
- â–¶ YouTube: youtube.com/@Academia_SD
- 💬 Discord: discord.gg/Syuaduy678
- ☕ Ko-Fi: ko-fi.com/academiasd
- Downloads last month
- 35
Model tree for AcademiaSD/Z-Image-NF4-for-LoRA-Training
Base model
Tongyi-MAI/Z-Image