FLUX.1-dev ConvRot for ComfyUI
Native ComfyUI ConvRot quantizations of FLUX.1-dev, with three diffusion-model choices and an optional INT8 ConvRot T5-XXL text encoder.
These are post-training quantized derivatives of black-forest-labs/FLUX.1-dev at source revision 3de623fc3c33e44ffbe2bad470d0f45bccf2eb21. They load with ComfyUI's core UNETLoader; no custom INT4 loader is required.
License: The model weights are subject to the upstream FLUX.1-dev Non-Commercial License. Access and use require compliance with that license and the upstream acceptable-use terms. See License before downloading.
Files and recommended combinations
| File | Role | Size | Recommendation |
|---|---|---|---|
FLUX.1-dev-w8a8-convrot.safetensors |
Whole-DiT ConvRot W8A8 | 11.981 GB | Default DiT: best storage/fidelity balance |
FLUX.1-dev-int8_convrot.safetensors |
Partial ConvRot INT8 with sensitive tensors preserved | 15.920 GB | Quality-first DiT: closest tested quantized variant to BF16 |
FLUX.1-dev-paper-w4a4-convrot.safetensors |
Paper-aligned whole-DiT ConvRot W4A4 | 6.248 GB | Memory-first / experimental |
t5xxl_flux1_int8_convrot.safetensors |
Experimental T5-XXL ConvRot W8A8 | 4.898 GB | Optional text-encoder companion |
The four uploadable files total 39.047 GB. The Original BF16 model shown in the comparison is a reference and is not included.
Recommended pipelines:
- Lowest combined storage:
FLUX.1-dev-w8a8-convrot.safetensors+t5xxl_flux1_int8_convrot.safetensors. - Default quality/storage balance:
FLUX.1-dev-w8a8-convrot.safetensors+ upstream BF16 T5-XXL. - Quality-first:
FLUX.1-dev-int8_convrot.safetensors+ upstream BF16 T5-XXL. - Memory-first:
FLUX.1-dev-paper-w4a4-convrot.safetensors+ upstream BF16 T5-XXL.
The optional INT8 T5 does not contain CLIP-L. FLUX still requires CLIP-L and a VAE.
What W8A8 and W4A4 mean
ConvRot stores rotated quantized weights plus per-layer metadata. At runtime, the quantized weight layout dispatches the matching ComfyUI/comfy-kitchen linear operator. The current BF16 activation is rotated and dynamically quantized for the matrix multiplication; activations are not stored in the checkpoint.
- W8A8: INT8 weights and dynamically quantized INT8 activations for selected linear layers.
- W4A4: packed INT4 weights and dynamically quantized INT4 activations for selected linear layers.
- Biases, norms, embeddings, and tensors outside each policy remain at their declared higher precision.
Rotation group size is 256. The W4A4 artifact uses quantization group size 64.
ComfyUI installation
Place the selected files as follows:
ComfyUI/models/diffusion_models/FLUX.1-dev-w8a8-convrot.safetensors
# or FLUX.1-dev-int8_convrot.safetensors
# or FLUX.1-dev-paper-w4a4-convrot.safetensors
ComfyUI/models/text_encoders/t5xxl_flux1_int8_convrot.safetensors # optional
You must separately obtain the standard FLUX assets under their applicable licenses:
ComfyUI/models/text_encoders/clip_l_flux1.safetensors
ComfyUI/models/text_encoders/t5xxl_flux1_bf16.safetensors # unless using the INT8 T5 here
ComfyUI/models/vae/ae_flux1.safetensors
Use a current ComfyUI and comfy-kitchen build with native ConvRot support. The validated environment was:
- ComfyUI
0.28.0 - comfy-kitchen
0.2.20 - comfy-cli
1.11.1 - NVIDIA CUDA GPU
W4A4 requires SM 8.0 or newer. The first run may compile kernels and take longer.
Load the diffusion model with core UNETLoader and weight_dtype=default. Load CLIP-L and the chosen T5 with DualCLIPLoader(type=flux).
A ready-to-use API/comfy-cli workflow is provided at
workflows/FLUX1_Dev_ConvRot_API.json.
It defaults to Whole W8A8 + INT8 T5. Change the diffusion-model and/or T5
filename to test another combination.
A maintained copy is available in the general
comfy-workflows
repository for direct API use.
Paired quality evaluation
Twenty 1024×1024 images were generated per pipeline with identical prompts, seeds, sampler, scheduler, guidance, steps, CLIP-L, and VAE. Metrics compare final 8-bit RGB images to Original BF16.
| Pipeline | Global PSNR vs BF16 | Mean per-pair PSNR | Low-VRAM maximum peak |
|---|---|---|---|
| Whole W8A8 + INT8 T5 | 26.559 dB | 28.872 dB | 16,265 MiB |
| Whole W8A8 + BF16 T5 | 27.399 dB | 29.449 dB | 16,298 MiB |
| Partial INT8 + BF16 T5 | 27.888 dB | 29.857 dB | 20,351 MiB |
| Paper W4A4 + BF16 T5 | 17.926 dB | 19.041 dB | 10,797 MiB |
The isolated T5 comparison—same Whole W8A8 DiT, BF16 T5 versus INT8 T5—measured 27.389 dB global PSNR and 30.054 dB mean per-pair PSNR.
PSNR measures paired pixel similarity, not aesthetics, prompt adherence, anatomy, typography, or human preference. Small spatial changes can sharply reduce PSNR. LPIPS and human-preference metrics were not computed.
Full records:
VRAM methodology
Each benchmark cell used a fresh ComfyUI process. Total physical GPU memory was sampled every 20 ms under an exclusive GPU lock. Settings were 1024×1024, batch 1, 20 steps, Euler/simple, CFG 1.0, FLUX guidance 3.5. --lowvram is an offloading result; hardware, runtime versions, attention backend, and workflow can materially change memory use.
The no-offloading --highvram maximum peaks for the four publication pipelines were 19,998 MiB, 21,342 MiB, 21,342 MiB, and 18,974 MiB in the table's order. Original BF16 hit a capacity failure in that 24 GiB environment.
Quantization policies
- Whole W8A8: all 304 paper-scoped 2-D transformer-block linear weights use tensorwise INT8 ConvRot; 476 tensors remain BF16 or otherwise preserved.
- Partial INT8: profile-based selective INT8 ConvRot with a larger BF16-preserved set. This is the largest and highest-PSNR tested quantized DiT.
- Paper W4A4: 301 paper-scoped transformer-block linears use W4A4; 479 tensors remain BF16 or otherwise preserved. ComfyUI's fused
single_blocks.*.linear1expands some paper component-level decisions to the fused tensor. - INT8 T5-XXL: 168 encoder linear projections use W8A8; 51 embedding, relative-attention-bias, and normalization tensors remain BF16. This T5 policy is experimental and was not evaluated by the ConvRot paper.
See manifests/conversion_summaries.json and release_manifest.json for artifact-level provenance and hashes.
Limitations
- These are post-training quantizations, not fine-tuned models.
- The W4A4 result diverges substantially more from BF16 and is labeled experimental.
- INT8 T5 image similarity does not establish long-prompt comprehension equivalence.
- Results cover 20 fixed prompts and one validated runtime/hardware environment.
- No safety classifier or content filter is included.
- Quantized checkpoints do not remove the limitations or usage restrictions of FLUX.1-dev.
License
The quantized weights are derivatives of FLUX.1-dev and are provided under the FLUX.1-dev Non-Commercial License. Review and comply with the exact upstream license:
- License: https://huggingface.co/black-forest-labs/FLUX.1-dev/blob/3de623fc3c33e44ffbe2bad470d0f45bccf2eb21/LICENSE.md
- Base model: https://huggingface.co/black-forest-labs/FLUX.1-dev
The helper workflow/scripts do not grant additional rights to the model weights. See LICENSE_NOTICE.md.
Credits
- Base model: Black Forest Labs,
FLUX.1-dev - Quantized runtime: ComfyUI and comfy-kitchen
- Conversion tooling: Starnodes Model Converter
- Method reference: ConvRot: Rotation-Based Plug-and-Play 4-bit Quantization for Diffusion Transformers, arXiv:2512.03673
Model tree for SearchingMan/FLUX.1-dev-ConvRot
Base model
black-forest-labs/FLUX.1-dev