Mage-Flow-NVFP4-AJH

Mage-Flow-NVFP4-AJH is a portable, runnable NVFP4 package for microsoft/Mage-Flow. It combines a native NVFP4 Mage transformer with a complete mixed NVFP4/FP8 Qwen3-VL text encoder.

ComfyUI

Ready-to-use custom nodes are available at AJH-Code/ComfyUI-MageFlow-NVFP4-AJH. The plugin provides loader and generation nodes, downloads this complete model repository, and returns a standard ComfyUI IMAGE. Its initial release targets Linux x86-64, Python 3.11, and NVIDIA Blackwell SM120 GPUs; follow the exact runtime requirements documented in the plugin repository.

Showcase

Cyborg woman gazing toward a star-filled sky, generated by Mage-Flow-NVFP4-AJH

Generated directly with the released full NVFP4 package, without upscaling or post-processing.

A 4K resolution high detail photo realistic image of the top half of a cyborg woman with dark black hair, striking blue eyes that have a very subtle glow in the iris, standing side profile, head tilted up towards the sky with a questioning expression, she has subtle gaps in her skin that hint at a robotic nature, outdoor forest night setting, sky filled with bright brilliant stars that glow against the dark setting, nebula visible

Settings: 1280×1280, 20 steps, CFG 5, static shift 6, seed 3334072683.

This is a complete Hugging Face component-layout repository, not an overlay:

model_index.json
transformer/
  config.json
  diffusion_pytorch_model-00001-of-00004.safetensors
  diffusion_pytorch_model.safetensors.index.json
text_encoder/
  config.json
  model.safetensors
vae/
scheduler/

The transformer shards include every retained BF16 tensor alongside the NVFP4 module state. The VAE, scheduler, text-encoder configuration, tokenizer, and processor are also included. Running the downloaded repository does not fetch BF16 weights from the base model.

Quantization policy

Mage transformer:

  • 48 image/text MLP up/down projections use resident W4A4 NVFP4.
  • Attention, modulation, normalization, and all non-target tensors remain BF16.
  • The 48 original BF16 target weight matrices are absent. Their biases remain BF16 and are stored normally inside the quantized modules.
  • The complete transformer is stored as four indexed standard Safetensors shards. Packed NVFP4 buffers use their module state-dict names: packed_weight, weight_scales, weight_scale, and bias.

Qwen3-VL text encoder:

  • Blocks 2–33: 224 NVFP4 projections.
  • Blocks 1 and 34: 14 FP8 projections.
  • Blocks 0 and 35, embeddings, norms, biases, and the vision tower remain BF16.
  • Packaged text checkpoint size: 4,031,376,064 bytes.
  • The packaged loader constructs the model directly from this file: 475 non-quantized tensors and 238 packed projections, with no unresolved meta tensors.

Requirements

The prebuilt runtime was tested on:

  • NVIDIA RTX 50-series / SM120
  • Linux x86-64
  • CUDA 13.1
  • Python 3.11
  • PyTorch 2.13.0+cu130
  • comfy-kitchen==0.2.22
  • flash-attn==2.8.3

Create a local environment:

python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
CUDA_HOME=/usr/local/cuda-13.1 \
  python -m pip install --no-build-isolation flash-attn==2.8.3

The included binaries are for the exact tested stack. Rebuild them after changing PyTorch, CUDA, or the C++ ABI:

CUDA_HOME=/usr/local/cuda-13.1 \
PYTHON_BIN="$PWD/.venv/bin/python" \
  ./build_native.sh

Generate an image

Expose exactly one SM120 GPU:

CUDA_VISIBLE_DEVICES=0 .venv/bin/python generate.py \
  --prompt 'A detailed watercolor fox reading under an old oak tree' \
  --output fox.png \
  --height 1024 \
  --width 1024 \
  --steps 20 \
  --seed 1

When running from another copy of the scripts, --model also accepts the Hub repository id or a downloaded standard-layout directory:

CUDA_VISIBLE_DEVICES=0 .venv/bin/python generate.py \
  --model ajh-code/Mage-Flow-NVFP4-AJH \
  --prompt 'A lighthouse poster reading "ARCTIC LOOP"' \
  --output lighthouse.png

The command refuses to overwrite an existing output and writes a companion JSON report containing coverage, memory, timing, and environment information.

Measured results

Transformer:

  • Loaded packed transformer allocation: 5,628,438,016 bytes.
  • Measured transformer allocation saving: 2,604,638,208 bytes (2.4258 GiB).
  • Target BF16 source reads: 0.
  • Representative real up-projection: 3.26x eager and 3.25x compiled speedup over BF16.

Text encoder:

  • BF16 language-stack residency: 7.5453 GiB.
  • Packed mixed residency: 3.0427 GiB.
  • Measured saving: 4.5026 GiB.

Combined validation on an RTX 5060 Ti 16 GB:

  • All seven coverage, source-access, allocation, finite, pixel, latent, and text-policy gates passed.
  • Peak text stage with the packed transformer still resident: 10,437,891,584 bytes (9.72 GiB).
  • VAL-07 pixel cosine/NRMSE versus frozen BF16: 0.9927827428 / 0.1200022063.
  • VAL-07 latent cosine/NRMSE versus frozen BF16: 0.9339298065 / 0.3618087155.
  • The requested strings ARCTIC LOOP and NORTHERN COAST were both rendered correctly.

Combined native NVFP4 VAL-07 output

The standard-layout packaged loader is validated independently of the research-tree loader. It installs 48 native Mage projections and 238 packed Qwen projections and rejects any checkpoint containing the replaced BF16 transformer targets.

Quality and speed caveats

The mixed text encoder is functional and produced visually strong downstream images, but it does not meet our unusually strict embedding-similarity gate: mean token/pooled cosine was 0.9531366898 / 0.9751400001. All ten content screening verdicts and category lists remained unchanged.

The ten-case text workload was slower with the packed runtime (30.36 s) than BF16 (18.29 s). This package therefore claims major text-encoder VRAM savings, not a text-encoding speedup. Transformer projections are materially faster, but repeated matched end-to-end timing has not been completed.

Additional limitations:

  • Native execution is currently SM120-only.
  • The repository follows the Hugging Face component and Safetensors layout, but the new mage_flow_nvfp4 runtime is not yet built into stock Diffusers. Use the included loader.
  • CUDA graph compatibility is not claimed.
  • Generation and text-to-image are tested; Base, Turbo, and editing variants are not.
  • A controlled photorealistic 1024×1024 comparison found that the current NVFP4 transformer can produce fuzzy hair or skin microtexture for some seeds. Using the original BF16 text encoder did not remove it; the transformer quantization is the primary associated variable.
  • Four transformer-only held-out cases and one combined held-out case have been evaluated. This is not a broad benchmark.

Validate the download

MANIFEST.json records every distributed file except itself:

.venv/bin/python validate_release.py

Hashing the transformer shards and text checkpoint can take a little while.

License and attribution

  • Mage-Flow and the vendored Mage inference source are Copyright (c) 2026 Microsoft and licensed under MIT. See LICENSE and licenses/MAGE-MIT.txt.
  • Qwen3-VL and the mixed NVFP4/FP8 text checkpoint are licensed under Apache-2.0. See licenses/QWEN-APACHE-2.0.txt.
  • The text checkpoint was produced by InsecureErasure/Qwen3-VL-4B-Instruct-NVFP4 using learned rounding and comfy-kitchen.
Downloads last month
67
Safetensors
Model size
3B params
Tensor type
F32
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ajh-code/Mage-Flow-NVFP4-AJH

Quantized
(4)
this model