You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

By clicking "Agree", you agree to the FLUX Non-Commercial License Agreement and acknowledge the Acceptable Use Policy of Black Forest Labs.

Log in or Sign Up to review the conditions and access this model content.

FLUX.2 [klein] 9B for the Snapdragon NPU (PulseX Image)

This FLUX Model is licensed by Black Forest Labs Inc. under the FLUX Non-Commercial License. Copyright Black Forest Labs Inc. IN NO EVENT SHALL BLACK FOREST LABS INC. BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH USE OF THIS MODEL.

These files are a modified version of FLUX.2 [klein] 9B: quantised to 8 and 4 bits, laid out for the Hexagon NPU, and with the picture decoder compiled for it. They are not an official product of Black Forest Labs and are not endorsed, approved or validated by Black Forest Labs.

A copy of the license is in LICENSE.md. Any rights to use these files are granted to you directly by Black Forest Labs Inc. under that license: non-commercial use only.

What this is for

PulseX Image - a Windows program that makes pictures and posters on the NPU of a Snapdragon X laptop, without a graphics card: 512 x 512 in 22 seconds, about 0.15 Wh per picture for the whole computer. The files here are the model in the form that program runs. They are of no use to other programs: the weights are in .hexpack files in the NPU's own layout, and the .gguf beside each holds only the metadata and the few tensors the processor needs.

hf download Pexqman/FLUX.2-klein-9B-hexpack --local-dir models

Add --include "q8/*" "vae-qnn/*" for the Best quality only (17 GB) or --include "q4/*" "vae-qnn/*" for Faster (10 GB); everything is 26 GB. Then follow the README of PulseX Image.

Files

Folder Files Size What
q8/ flux-2-klein-9b-q8_0-slim.gguf + .hexpack 10.0 GB the transformer, Q8_0
q8/ qwen3-8b-flux2-q8_0-slim.gguf + .hexpack 6.2 GB the text encoder (Qwen3 8B as shipped with FLUX.2 [klein]; the 27 layers the model reads), Q8_0
q4/ flux-2-klein-9b-q4_0-slim.gguf + .hexpack 5.6 GB the transformer, Q4_0
q4/ qwen3-8b-flux2-q4_0-slim.gguf + .hexpack 3.3 GB the text encoder, Q4_0
vae-qnn/ flux2-vae-decoder-<W>x<H>_ctx_qnn.bin (512x512, 512x768, 768x512, 1024x1024) 0.5 GB the VAE decoder as a QNN context binary, one per picture size
vae-qnn/ flux2-vae-encoder-<W>x<H>_ctx_qnn.bin (512x512, 512x768, 768x512), bn.f32 0.2 GB the VAE encoder (for starting from a picture of your own) and the latents' statistics
sr-qnn/ quicksrnetlarge_512x512_ctx_qnn.bin 1.3 MB the upscaler for "Enlarge" - not FLUX: QuickSRNet-Large from Qualcomm AI Hub, BSD-3-Clause, see sr-qnn/LICENSE-QuickSRNet.txt

A .gguf and the .hexpack beside it belong together (a key stored in both) and must be kept in the same folder under the same name.

The FLUX Non-Commercial License covers the files in q8/, q4/ and vae-qnn/. The upscaler in sr-qnn/ is a separate work under its own license (BSD-3-Clause, Qualcomm's QuickSRNet-Large, compiled for 512 x 512 on the same chip); it is here so that one download gives everything the program needs.

Made how

  • From the bf16 weights of black-forest-labs/FLUX.2-klein-9B, with ggml's Q8_0 and Q4_0 quantisation of every block matrix; the modulation, embedding and output layers the processor computes stay as they were.
  • .hexpack: the same numbers, in the tiled layout the Hexagon NPU's matrix unit reads directly (pack tag hexagon-tiled-v2 of the PulseX fork of ggml-hexagon), so a block is read from disk and used without conversion.
  • The VAE: exported to ONNX per picture size and compiled to a QNN context binary (fp16) with Qualcomm's AI runtime 2.46, on a Snapdragon X Plus (X1P-42-100, Hexagon NPU v73). On other Snapdragon X chips these files are untested. They need QAIRT 2.46 or newer to load.

How close to the original

The whole chain (text encoder + transformer, 4 steps, 512 x 512) against the same chain in bf16 on the processor, cosine similarity of the final latents for three prompts:

prompt 1 prompt 2 prompt 3
Q8_0 0.990 0.993 0.949
Q4_0 0.938 0.935 0.826

Q8_0 gives the original's picture. Q4_0 gives a picture of the same quality that is another sample - composition and details move. Both take the same time per step on the NPU; Q4_0 is faster overall because less is read from disk.

Hardware

Snapdragon X Plus, 16 GB of memory, Windows 11 on ARM64. The NPU module of PulseX Image is self-signed: Windows loads it only on a machine set up to accept self-signed NPU modules (see the program's README).

Downloads last month
2
GGUF
Model size
0.4B params
Architecture
flux2
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Pexqman/FLUX.2-klein-9B-hexpack

Quantized
(54)
this model