Qwen-Image-2.1 on MLX, faster and exact

I wrote this renderer to run Qwen-Image-2.1 on a Mac mini M4 with 16 GB. The default path produces the same pixels as the stock MLX implementation, bit for bit, and needs about 10 % less time at 1024 × 1024. With long prompts the gap grows to about 1.9× per step, because the text tokens are computed once per image and reused at every later step.

The repository contains code only. The weights are the unmodified 4-bit mlx-community/Qwen-Image-2.1-MLX-4bit checkpoint, and the transformer and VAE modules come from mflux 0.20.

Results on a Mac mini M4 (16 GB)

1024 × 1024, 25 steps, 407-token prompt, all rows on MLX 0.32.2:

Path Time Output
mflux 0.20 transformer, same checkpoint (starting point) 594 s reference
This renderer, default 536 s pixel-identical
With Fast mode (first-block cache, threshold 0.08) 341 s approximate
Turbo, 6 steps with the Viggle LoRA 139 to 145 s approximate

Corrected on 2026-09-27. An earlier version of this page gave 619 s for the starting point. That run used MLX 0.32.0, while the other rows used MLX 0.32.2. Re-measured on 0.32.2 with the same prompt and seed, the starting point takes 594 s. About 30 % of the gap I first claimed came from the MLX upgrade, not from this code.

Median time per denoising step, on MLX 0.32.2. The latents before and after are identical:

Case Before After
384 × 384, 22-token prompt 2.68 s 2.50 s
384 × 384, 494-token prompt 4.82 s 2.55 s
1024 × 1024, 407-token prompt 23.57 s 21.27 s

At 384 px with a 494-token prompt, text makes up 46 % of the sequence. Without reuse, it costs almost half of every step. With reuse, the 494-token prompt runs almost as fast as the 22-token one.

Against ComfyUI on the same machine

Same M4, 1024 × 1024, 25 steps. ComfyUI runs qwen_image_2.1-Q4_K_M.gguf through ComfyUI-GGUF, using PyTorch on MPS.

Setup Time Relative
ComfyUI, GGUF Q4_K_M 736 s 1×
This renderer, default 536 s about 1.3 to 1.4×
This renderer, Fast mode 341 s about 2.1×
This renderer, Turbo 6 steps 139 to 145 s about 5×

The ComfyUI figure covers the whole job, text encoding included. My figures start after text encoding, which ComfyUI also does for this renderer. The two runs used different prompts. Given both differences, I quote a range and not a single ratio.

Original (594 s), this renderer (536 s) and Fast mode (341 s) at 1024

From left: mflux path, this renderer (identical pixels), Fast mode.

Base model at 25 steps against Turbo at 6 steps, 1024

Base model at 25 steps on the left. Turbo at 6 steps with the same seed in the middle, and with a second seed on the right.

Where the time goes, and what I changed

At 384 px with a 22-token prompt, one step pushes 598 tokens through 32 blocks. When I profiled it on MLX 0.32.0, that was about 8.5 TFLOP in 2.78 s, or 3.07 TFLOPS effective. The 4-bit matmuls took 81.5 ms per block, which is 94 % of the step. That leaves little to gain inside a step, so most of the savings come from doing less work.

Text K/V reuse (exact, on by default)

Qwen-Image-2.1 is a single-stream model, yet its text rows never see the image. They use the t = 0 modulation row and causal attention over text only. Their keys and values are therefore constant for the whole denoising run.

On step 1, I run mflux's unmodified joint forward and wrap the module-level scaled_dot_product_attention in qwen21_attention.py. The wrapper records the K/V of the causal (text) calls. Every later step processes image rows only and reads the stored text K/V.

I capture from the joint pass on purpose. A separate 22-row text-only pass rounds differently from the same rows inside a 598-row matmul, while 576, 598 and 1070 rows all give identical results, so MLX most likely switches quantized matmul kernels for small inputs. proj_out has the same row-count sensitivity, so it runs padded to the joint row count.

The cache costs 2 × 32 × tokens × 4096 × 2 bytes, about 0.5 MB per token (259 MB at 494 tokens). With CFG enabled, the negative prompt's K/V sit in a second slot. A 384 px CFG render with reuse matches a full recomputation to the last pixel.

Fused RoPE kernel (exact)

mflux applies rotary embeddings in fp32 with reshapes, slices and a stack. I replaced that with a single mx.fast.metal_kernel, which cuts each call from 8.64 ms to 0.99 ms at 1024. Matching MLX's output bit for bit took precise::fma(a, b, 0) for the products. Left alone, the Metal compiler contracts multiply-adds into FMAs, which changed about 100 of 4.4 million values. At install time the kernel checks itself against mflux's version and falls back to it on any mismatch.

Smaller exact changes

  • The denoising loop syncs with the host less often.
  • Several seeds of the same prompt share one model load and one text encoding. Small images are decoded as soon as each one is sampled.
  • Encoded prompts are cached. The key covers the ComfyUI workflow and the identity of the text encoder file.
  • At large sizes the transformer is released before the VAE loads. A guard fails the render if more than 1 GB of MLX memory is still held after the release, so leaks surface instead of pushing the Mac into swap.
  • Two other guards stop a render if free memory falls below 2 GB or if swap grows by more than 1 GB.

Optional modes (approximate or slower)

  • Fast mode is a first-block cache. Block 0 always runs, and blocks 1 to 31 are skipped while block 0's residual changes less than the threshold.
  • Turbo loads the Viggle turbo LoRA at runtime without merging it (Viggle reports that merging into bf16 loses about 30 % of the update). It uses the LoRA's fixed schedule of 5 to 8 steps.
  • A DPM++ 2M sampler is also available.
  • True CFG uses the negative prompt the way mflux does, noise = neg + g · (pos - neg). It needs a second forward pass per step, so it takes about twice as long.
  • Live previews come straight from the latents through a least-squares projection from 64 channels to RGB, without running the VAE.

Files

The weights and the model architecture are unchanged. Everything here loads mflux's modules, patches them and runs the sampling loop.

File Purpose
prefix_reuse.py Transformer subclass: joint step 1 with K/V capture, image-only steps after it, a second slot for CFG
fused_rope.py Metal RoPE kernel with the install-time self-check
mlx_renderer.py Denoising loop, batching, memory staging, CFG, previews and guards
mlx_loader.py Loads the 4-bit checkpoint into quantized layers and validates every weight key and config value
samplers.py DPM++ 2M and the turbo LoRA's fixed-node schedule
lora.py, turbo.py Unmerged runtime LoRA and the Turbo presets
conditioning_cache.py, conditioning_io.py, comfy_*.py Text encoding through ComfyUI, the prompt cache, shape and mask checks
latent_preview.py The fitted 64-channel-to-RGB preview map
benchmark.py Memory snapshots (free RAM, swap, process and MLX memory) for the guards and benchmarks

Quick install

hf download tillknuesting/qwen-image-2.1-mlx-fast --local-dir qwen-image-2.1-mlx-fast
cd qwen-image-2.1-mlx-fast
bash install.sh                  # add --install-comfyui if you do not have ComfyUI yet, --turbo for the Turbo LoRA
./render.sh "A ceramic teapot on a wooden table, soft window light"

install.sh checks for Apple Silicon and installs uv if it is missing. It then creates the Python environment, downloads the MLX weights (about 5 GB) and sets up ComfyUI for text encoding: both custom nodes, their requirements, and the Qwen3-VL text encoder (about 5 GB). It expects ComfyUI in ~/ComfyUI; use --comfy-dir for another location. Running it again skips everything that is already in place.

render.sh starts ComfyUI when it is not running and saves the picture in out/. A bare prompt renders at 1024 × 1024 with 25 steps and a random seed. Any option of run_mlx_job.py works too, for example ./render.sh "…" --turbo or ./render.sh --prompt "…" --width 832 --height 1216 --seed 42 --output me.png.

I tested both scripts on the M4 against an existing ComfyUI, without the two large downloads. That run covered the checks, the environment, custom-node detection, the settings file, and renders through render.sh.

Manual setup

You need an Apple Silicon Mac and uv, which installs Python 3.13 for you. 16 GB of memory is enough for 1024 × 1024.

  1. Get the code and its dependencies:

    hf download tillknuesting/qwen-image-2.1-mlx-fast --local-dir qwen-image-2.1-mlx-fast
    cd qwen-image-2.1-mlx-fast
    uv sync
    
  2. Download the transformer and VAE from mlx-community (about 5 GB, into models/):

    uv run python scripts/download_mlx_components.py
    
  3. Set up text encoding. The prompt is still encoded by ComfyUI, so you need a recent ComfyUI with Qwen-Image-2.1 support and two custom nodes. ComfyUI-GGUF loads GGUF models, and ComfyUI-GGUF-Qwen3VL-TE patches it so it accepts a Qwen3-VL text encoder; without the patch, encoding fails with a normalized_shape error. Install ComfyUI-GGUF's own requirements into ComfyUI's Python environment.

    cd /path/to/ComfyUI/custom_nodes
    git clone https://github.com/city96/ComfyUI-GGUF
    git clone https://github.com/pottokao-dotcom/ComfyUI-GGUF-Qwen3VL-TE
    hf download Qwen/Qwen3-VL-8B-Instruct-GGUF Qwen3VL-8B-Instruct-Q4_K_M.gguf \
      --local-dir /path/to/ComfyUI/models/text_encoders
    

    Then start ComfyUI and leave it running. My script sends it the prompt, reads back the saved embeddings and asks it to free its memory before MLX starts rendering. Porting the encoder to MLX would remove this step, and it is next on my list.

  4. Optional, for Turbo: download the Viggle LoRA. Its licence (qwen-research) does not allow me to redistribute it.

    hf download Viggle/Qwen-Image-2.1-viggle-turbo Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r128.safetensors \
      --local-dir models/loras/viggle-turbo
    

Run

uv run python scripts/run_mlx_job.py \
  --clip-name Qwen3VL-8B-Instruct-Q4_K_M.gguf \
  --prompt "Candid photograph of an elderly man reading at a kitchen table, soft window light" \
  --width 1024 --height 1024 --steps 25 --seed 42 \
  --output out/man.png

The script expects ComfyUI at http://127.0.0.1:8188 with its output folder at ~/ComfyUI/output. If yours differ, pass --comfy-url and --comfy-output /path/to/ComfyUI/output. The first render of a new prompt includes text encoding. A repeated prompt comes from the cache in artifacts/conditioning_cache.

On the M4 I checked a fresh uv sync, the test suite and a render with this script, using weights I had already downloaded and a Q4_K_M Qwen3-VL 8B encoder of my own. A 384 × 384 render with 4 steps took 39 s, text encoding included. I have not yet run the official Qwen3VL-8B-Instruct GGUF through this setup.

Options I use most:

Option Effect
--batch-file jobs.json Renders a list of {seed, steps, output} entries with one model load
--first-block-cache 0.08 Fast mode
--turbo --turbo-steps 6 Turbo (needs the LoRA from step 4)
--negative-prompt "…" --cfg 2.5 True CFG
--live-preview --progress-file progress.jsonl Per-step events with memory figures, plus preview images
--no-reuse-prefix Turns text reuse off, for checking exactness yourself

scripts/render_conditioning.py renders from exported embeddings without ComfyUI. scripts/check_prefix_exactness.py compares the reuse path with the joint path.

Tests

uv run pytest

Licences and credits

My code is MIT-licensed. Qwen-Image-2.1 belongs to the Alibaba Qwen team and has its own licence. The MLX conversion is by mlx-community, and the transformer and VAE implementation is from mflux (MIT). The Viggle turbo LoRA uses the non-commercial qwen-research licence, so it is not included here; download it from Viggle's repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support