FLUX.2 [klein] 4B - bitsandbytes NF4

black-forest-labs/FLUX.2-klein-4B with the transformer and the Qwen3 text encoder quantised to 4-bit NF4 (double quantisation, bf16 compute) using bitsandbytes, saved as a standard diffusers pipeline. Same weights as the original at about a quarter of the size, for consumer NVIDIA GPUs.

import torch
from diffusers import Flux2KleinPipeline

pipe = Flux2KleinPipeline.from_pretrained("hlhc/FLUX.2-klein-4B-bnb-4bit", torch_dtype=torch.bfloat16).to("cuda")
out = pipe(prompt="...", image=image, num_inference_steps=4, guidance_scale=1.0).images[0]

Needs diffusers>=0.36, transformers, accelerate, bitsandbytes>=0.45 and a CUDA GPU.

Provenance

Quantised from black-forest-labs/FLUX.2-klein-4B with bitsandbytes NF4 (double quantisation), compute dtype bf16.

Smoke test on NVIDIA L4, whole pipeline on the GPU: 4.8 s for a 768 x 512 edit in 4 steps, peak VRAM 5.8 GB. Input and output: smoke_test_input.png, smoke_test_output.png. Cards with less memory than the peak should use pipe.enable_model_cpu_offload().

Licence: Apache-2.0, unchanged from Black Forest Labs (see LICENSE.md).

Downloads last month
114
Safetensors
Model size
2B params
Tensor type
F32
路
BF16
路
U8
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for hlhc/FLUX.2-klein-4B-bnb-4bit

Quantized
(51)
this model