Instructions to use tillknuesting/qwen-image-2.1-mlx-fast with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use tillknuesting/qwen-image-2.1-mlx-fast with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir qwen-image-2.1-mlx-fast tillknuesting/qwen-image-2.1-mlx-fast
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Qwen-Image-2.1 on MLX, faster and exact
I wrote this renderer to run Qwen-Image-2.1 on a Mac mini M4 with 16 GB. The default path produces the same pixels as the stock MLX implementation, bit for bit, and needs about 10 % less time at 1024 × 1024. With long prompts the gap grows to about 1.9× per step, because the text tokens are computed once per image and reused at every later step.
The repository contains code only. The weights are the unmodified 4-bit
mlx-community/Qwen-Image-2.1-MLX-4bit checkpoint, and
the transformer and VAE modules come from mflux 0.20.
Results on a Mac mini M4 (16 GB)
1024 × 1024, 25 steps, 407-token prompt, all rows on MLX 0.32.2:
| Path | Time | Output |
|---|---|---|
| mflux 0.20 transformer, same checkpoint (starting point) | 594 s | reference |
| This renderer, default | 536 s | pixel-identical |
| With Fast mode (first-block cache, threshold 0.08) | 341 s | approximate |
| Turbo, 6 steps with the Viggle LoRA | 139 to 145 s | approximate |
Corrected on 2026-09-27. An earlier version of this page gave 619 s for the starting point. That run used MLX 0.32.0, while the other rows used MLX 0.32.2. Re-measured on 0.32.2 with the same prompt and seed, the starting point takes 594 s. About 30 % of the gap I first claimed came from the MLX upgrade, not from this code.
Median time per denoising step, on MLX 0.32.2. The latents before and after are identical:
| Case | Before | After |
|---|---|---|
| 384 × 384, 22-token prompt | 2.68 s | 2.50 s |
| 384 × 384, 494-token prompt | 4.82 s | 2.55 s |
| 1024 × 1024, 407-token prompt | 23.57 s | 21.27 s |
At 384 px with a 494-token prompt, text makes up 46 % of the sequence. Without reuse, it costs almost half of every step. With reuse, the 494-token prompt runs almost as fast as the 22-token one.
Against ComfyUI on the same machine
Same M4, 1024 × 1024, 25 steps. ComfyUI runs qwen_image_2.1-Q4_K_M.gguf through ComfyUI-GGUF, using PyTorch on MPS.
| Setup | Time | Relative |
|---|---|---|
| ComfyUI, GGUF Q4_K_M | 736 s | 1× |
| This renderer, default | 536 s | about 1.3 to 1.4× |
| This renderer, Fast mode | 341 s | about 2.1× |
| This renderer, Turbo 6 steps | 139 to 145 s | about 5× |
The ComfyUI figure covers the whole job, text encoding included. My figures start after text encoding, which ComfyUI also does for this renderer. The two runs used different prompts. Given both differences, I quote a range and not a single ratio.
From left: mflux path, this renderer (identical pixels), Fast mode.
Base model at 25 steps on the left. Turbo at 6 steps with the same seed in the middle, and with a second seed on the right.
Where the time goes, and what I changed
At 384 px with a 22-token prompt, one step pushes 598 tokens through 32 blocks. When I profiled it on MLX 0.32.0, that was about 8.5 TFLOP in 2.78 s, or 3.07 TFLOPS effective. The 4-bit matmuls took 81.5 ms per block, which is 94 % of the step. That leaves little to gain inside a step, so most of the savings come from doing less work.
Text K/V reuse (exact, on by default)
Qwen-Image-2.1 is a single-stream model, yet its text rows never see the image. They use the t = 0 modulation row and causal attention over text only. Their keys and values are therefore constant for the whole denoising run.
On step 1, I run mflux's unmodified joint forward and wrap the module-level scaled_dot_product_attention in
qwen21_attention.py. The wrapper records the K/V of the causal (text) calls. Every later step processes image rows
only and reads the stored text K/V.
I capture from the joint pass on purpose. A separate 22-row text-only pass rounds differently from the same rows
inside a 598-row matmul, while 576, 598 and 1070 rows all give identical results, so MLX most likely switches
quantized matmul kernels for small inputs. proj_out has the same row-count sensitivity, so it runs padded to the
joint row count.
The cache costs 2 × 32 × tokens × 4096 × 2 bytes, about 0.5 MB per token (259 MB at 494 tokens). With CFG enabled, the negative prompt's K/V sit in a second slot. A 384 px CFG render with reuse matches a full recomputation to the last pixel.
Fused RoPE kernel (exact)
mflux applies rotary embeddings in fp32 with reshapes, slices and a stack. I replaced that with a single
mx.fast.metal_kernel, which cuts each call from 8.64 ms to 0.99 ms at 1024. Matching MLX's output bit for bit took
precise::fma(a, b, 0) for the products. Left alone, the Metal compiler contracts multiply-adds into FMAs, which changed
about 100 of 4.4 million values. At install time the kernel checks itself against mflux's version and falls back to it
on any mismatch.
Smaller exact changes
- The denoising loop syncs with the host less often.
- Several seeds of the same prompt share one model load and one text encoding. Small images are decoded as soon as each one is sampled.
- Encoded prompts are cached. The key covers the ComfyUI workflow and the identity of the text encoder file.
- At large sizes the transformer is released before the VAE loads. A guard fails the render if more than 1 GB of MLX memory is still held after the release, so leaks surface instead of pushing the Mac into swap.
- Two other guards stop a render if free memory falls below 2 GB or if swap grows by more than 1 GB.
Optional modes (approximate or slower)
- Fast mode is a first-block cache. Block 0 always runs, and blocks 1 to 31 are skipped while block 0's residual changes less than the threshold.
- Turbo loads the Viggle turbo LoRA at runtime without merging it (Viggle reports that merging into bf16 loses about 30 % of the update). It uses the LoRA's fixed schedule of 5 to 8 steps.
- A DPM++ 2M sampler is also available.
- True CFG uses the negative prompt the way mflux does,
noise = neg + g · (pos - neg). It needs a second forward pass per step, so it takes about twice as long. - Live previews come straight from the latents through a least-squares projection from 64 channels to RGB, without running the VAE.
Files
The weights and the model architecture are unchanged. Everything here loads mflux's modules, patches them and runs the sampling loop.
| File | Purpose |
|---|---|
prefix_reuse.py |
Transformer subclass: joint step 1 with K/V capture, image-only steps after it, a second slot for CFG |
fused_rope.py |
Metal RoPE kernel with the install-time self-check |
mlx_renderer.py |
Denoising loop, batching, memory staging, CFG, previews and guards |
mlx_loader.py |
Loads the 4-bit checkpoint into quantized layers and validates every weight key and config value |
samplers.py |
DPM++ 2M and the turbo LoRA's fixed-node schedule |
lora.py, turbo.py |
Unmerged runtime LoRA and the Turbo presets |
conditioning_cache.py, conditioning_io.py, comfy_*.py |
Text encoding through ComfyUI, the prompt cache, shape and mask checks |
latent_preview.py |
The fitted 64-channel-to-RGB preview map |
benchmark.py |
Memory snapshots (free RAM, swap, process and MLX memory) for the guards and benchmarks |
Quick install
hf download tillknuesting/qwen-image-2.1-mlx-fast --local-dir qwen-image-2.1-mlx-fast
cd qwen-image-2.1-mlx-fast
bash install.sh # add --install-comfyui if you do not have ComfyUI yet, --turbo for the Turbo LoRA
./render.sh "A ceramic teapot on a wooden table, soft window light"
install.sh checks for Apple Silicon and installs uv if it is missing. It then creates the Python environment,
downloads the MLX weights (about 5 GB) and sets up ComfyUI for text encoding: both custom nodes, their requirements,
and the Qwen3-VL text encoder (about 5 GB). It expects ComfyUI in ~/ComfyUI; use --comfy-dir for another
location. Running it again skips everything that is already in place.
render.sh starts ComfyUI when it is not running and saves the picture in out/. A bare prompt renders at
1024 × 1024 with 25 steps and a random seed. Any option of run_mlx_job.py works too, for example
./render.sh "…" --turbo or ./render.sh --prompt "…" --width 832 --height 1216 --seed 42 --output me.png.
I tested both scripts on the M4 against an existing ComfyUI, without the two large downloads. That run covered the
checks, the environment, custom-node detection, the settings file, and renders through render.sh.
Manual setup
You need an Apple Silicon Mac and uv, which installs Python 3.13 for you. 16 GB of memory is enough for 1024 × 1024.
Get the code and its dependencies:
hf download tillknuesting/qwen-image-2.1-mlx-fast --local-dir qwen-image-2.1-mlx-fast cd qwen-image-2.1-mlx-fast uv syncDownload the transformer and VAE from mlx-community (about 5 GB, into
models/):uv run python scripts/download_mlx_components.pySet up text encoding. The prompt is still encoded by ComfyUI, so you need a recent ComfyUI with Qwen-Image-2.1 support and two custom nodes. ComfyUI-GGUF loads GGUF models, and ComfyUI-GGUF-Qwen3VL-TE patches it so it accepts a Qwen3-VL text encoder; without the patch, encoding fails with a
normalized_shapeerror. Install ComfyUI-GGUF's own requirements into ComfyUI's Python environment.cd /path/to/ComfyUI/custom_nodes git clone https://github.com/city96/ComfyUI-GGUF git clone https://github.com/pottokao-dotcom/ComfyUI-GGUF-Qwen3VL-TE hf download Qwen/Qwen3-VL-8B-Instruct-GGUF Qwen3VL-8B-Instruct-Q4_K_M.gguf \ --local-dir /path/to/ComfyUI/models/text_encodersThen start ComfyUI and leave it running. My script sends it the prompt, reads back the saved embeddings and asks it to free its memory before MLX starts rendering. Porting the encoder to MLX would remove this step, and it is next on my list.
Optional, for Turbo: download the Viggle LoRA. Its licence (
qwen-research) does not allow me to redistribute it.hf download Viggle/Qwen-Image-2.1-viggle-turbo Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r128.safetensors \ --local-dir models/loras/viggle-turbo
Run
uv run python scripts/run_mlx_job.py \
--clip-name Qwen3VL-8B-Instruct-Q4_K_M.gguf \
--prompt "Candid photograph of an elderly man reading at a kitchen table, soft window light" \
--width 1024 --height 1024 --steps 25 --seed 42 \
--output out/man.png
The script expects ComfyUI at http://127.0.0.1:8188 with its output folder at ~/ComfyUI/output. If yours differ,
pass --comfy-url and --comfy-output /path/to/ComfyUI/output. The first render of a new prompt includes text
encoding. A repeated prompt comes from the cache in artifacts/conditioning_cache.
On the M4 I checked a fresh uv sync, the test suite and a render with this script, using weights I had already
downloaded and a Q4_K_M Qwen3-VL 8B encoder of my own. A 384 × 384 render with 4 steps took 39 s, text encoding
included. I have not yet run the official Qwen3VL-8B-Instruct GGUF through this setup.
Options I use most:
| Option | Effect |
|---|---|
--batch-file jobs.json |
Renders a list of {seed, steps, output} entries with one model load |
--first-block-cache 0.08 |
Fast mode |
--turbo --turbo-steps 6 |
Turbo (needs the LoRA from step 4) |
--negative-prompt "…" --cfg 2.5 |
True CFG |
--live-preview --progress-file progress.jsonl |
Per-step events with memory figures, plus preview images |
--no-reuse-prefix |
Turns text reuse off, for checking exactness yourself |
scripts/render_conditioning.py renders from exported embeddings without ComfyUI. scripts/check_prefix_exactness.py
compares the reuse path with the joint path.
Tests
uv run pytest
Licences and credits
My code is MIT-licensed. Qwen-Image-2.1 belongs to the Alibaba Qwen team and has its own licence. The MLX conversion
is by mlx-community, and the transformer and VAE implementation is from mflux (MIT). The Viggle turbo LoRA uses the
non-commercial qwen-research licence, so it is not included here; download it from Viggle's repository.

