MiniMax-H3 Β· FastH3-NVFP4-rotated (single-file, 16 GB)
A rotated-NVFP4 quantization of FastH3 β FastVideo's 4-step, VSA-distilled MiniMax-H3 (video + synchronized audio) β as one self-contained ~12.8 GB file that runs fully resident on a single 16 GB Blackwell card, zero CPU offload, on the native nunchaku W4A4 kernel. This is the fast line (4-step, ~5Γ fewer steps). For the higher-quality 8-step line see MiniMax-H3-NVFP4-rotated.
Loads with our
H3RotNVFP4Loadernode β get it at H3-RotNVFP4-ComfyUI-Loader (why a custom node, full method, and dependency versions are there).
Measured (1Γ RTX 5070 Ti 16 GB, warm, DiT-only)
- Genuine 16 GB residency, zero offload β ~1.2 GB CPU offload (base-class, NOT the 26 GB an un-shrunk AdaLN needs), peak VRAM 13.7 GB.
- 480p (832Γ480): ~5.7 s/step Γ 4 β full 124-frame clip in ~25β30 s of denoise. Max resident clip @768p: 124 f / 5.17 s.
- Audio preserved (clean 32 kHz stereo). File 12.8 GB (T1, all-NVFP4 blocks + rank-32 curve-basis AdaLN + bf16 islands).
Sampler β use Euler (important)
Run with the euler sampler + simple schedule (sigma shift 12), CFG 1.0, 4 steps β the settings FastH3 was distilled for. This is not our quantization: it's a standard sampler choice, and it matters at 4 steps.
eulerβ coherent, temporally stable video (recommended default; the includedworkflow.jsonuses it).res_multistep(a 2nd-order solver) looks a touch sharper but strobes on high-contrast neon / rain / water β the composition churns frame-to-frame. Avoid it here.- Want more sharpness/motion on the stable Euler base? Raise steps (6β8) β trades some speed. For maximum sharpness, the 8-step base line is crisper still.
Pick sampler/steps per scene and taste β the model is the same either way.
Samples
![]() close-up face β eyelashes, iris, skin |
![]() forest waterfall β moss, spray |
![]() sailboat / sea β rigging, sun-glitter |
![]() sci-fi rooftop β flying cars, neon |
Full MP4 + synchronized audio in samples/. Rotated-NVFP4 vs fp8 vs bf16 quality comparison (frames + IQA table) in quality/ β same perceptual sharpness as fp8/bf16, but NVFP4 runs the native fp4 kernel (fp8 falls back to bf16 GEMM). (GIFs are video-only previews; use the euler sampler per above for coherent motion.)
Setup
- Install the loader node + nunchaku β see H3-RotNVFP4-ComfyUI-Loader (verified
torch 2.12.1+cu130+ a nunchaku cu13/torch2.12 wheel; Linux + Blackwell RTX 50-series / GB10). - Download
h3_fasth3_T1.safetensorsintoComfyUI/models/diffusion_models/. - Load
workflow.json: H3RotNVFP4Loader (ckpt_name= the file), text encoder on CPU (CLIPLoader device=cpu), 4 steps. Log marker:[rot] single-file wrapped 200 linears (200 rotation-NVFP4 + 0 bf16-protected).
Credits & license
MiniMax-H3 (MiniMaxAI) Β· FastVideo (FastH3 4-step + VSA distillation) Β· nunchaku / SVDQuant (NVFP4 W4A4 kernel) Β· ComfyUI. Derivative under the MiniMax-H3 Community License β commercial use with attribution; US/UK/EU/South Korea users apply to MiniMax; >$20M/yr needs separate authorization. Verify for your use.
Model tree for pottokao/MiniMax-H3-FastH3-NVFP4-rotated
Base model
MiniMaxAI/MiniMax-H3


