MiniMax H3 โ€” NVFP4 + FP8 (fl2va)

A single ComfyUI checkpoint for MiniMax H3: NVFP4 for the bulk MLP weights, FP8 (tensor-scale) for attn.qkv_proj, AdaLN modulation pruned to a lookup table. fl2va (first/last-frame) mode only. Requires a Blackwell GPU (RTX 50-series, B100/B200).

File Size What it is
minimax_h3_fl2va_pruned_nvfp4_fp8.safetensors 20 GB NVFP4 MLP + FP8 attn.qkv_proj, AdaLN-pruned

attn.qkv_proj is quantized fresh from the original BF16 weights (default ctq FP8 output) and spliced into the NVFP4 base. Tested and confirmed working in ComfyUI.

Looking for other variants?

See rockerBOO/minimax-h3-nvfp4-convrot for the full set โ€” ref2va support, the recommended INT8 ConvRot attn.qkv_proj variant (same size, faster), the unpruned base, experimental INT4 ConvRot variants, and full quantization-method documentation. This file is the FP8 speed-comparison baseline for that repo's ConvRot INT8 variant.

License

MiniMax H3 Community License Agreement, inherited from MiniMaxAI/MiniMax-H3 (repackaged by Comfy-Org/MiniMax-H3).

MiniMax H3 is licensed under the MiniMax H3 Community License Agreement, Copyright ยฉ 2026 MiniMax. All Rights Reserved.

Downloads last month
31,370
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for rockerBOO/minimax-h3-nvfp4-fp8

Quantized
(54)
this model