Instructions to use hlhc/FLUX.2-klein-4B-bnb-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use hlhc/FLUX.2-klein-4B-bnb-4bit with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("hlhc/FLUX.2-klein-4B-bnb-4bit", dtype=torch.bfloat16, device_map="cuda") prompt = "Turn this cat into a dog" input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png") image = pipe(image=input_image, prompt=prompt).images[0] - Notebooks
- Google Colab
- Kaggle
FLUX.2 [klein] 4B - bitsandbytes NF4
black-forest-labs/FLUX.2-klein-4B with the transformer and the Qwen3 text encoder quantised to 4-bit NF4
(double quantisation, bf16 compute) using bitsandbytes, saved as a standard diffusers pipeline.
Same weights as the original at about a quarter of the size, for consumer NVIDIA GPUs.
import torch
from diffusers import Flux2KleinPipeline
pipe = Flux2KleinPipeline.from_pretrained("hlhc/FLUX.2-klein-4B-bnb-4bit", torch_dtype=torch.bfloat16).to("cuda")
out = pipe(prompt="...", image=image, num_inference_steps=4, guidance_scale=1.0).images[0]
Needs diffusers>=0.36, transformers, accelerate, bitsandbytes>=0.45 and a CUDA GPU.
Provenance
Quantised from black-forest-labs/FLUX.2-klein-4B with bitsandbytes NF4 (double quantisation), compute dtype bf16.
Smoke test on NVIDIA L4, whole pipeline on the GPU: 4.8 s for a
768 x 512 edit in 4 steps, peak VRAM 5.8 GB. Input and output: smoke_test_input.png,
smoke_test_output.png. Cards with less memory than the peak should use
pipe.enable_model_cpu_offload().
Licence: Apache-2.0, unchanged from Black Forest Labs (see LICENSE.md).
- Downloads last month
- 114
Model tree for hlhc/FLUX.2-klein-4B-bnb-4bit
Base model
black-forest-labs/FLUX.2-klein-4B