FastH3 Trim, 8-step, NVFP4, for Blackwell GPUs and DGX Spark

FastH3 Trim is an experimental, smaller version of FastH3 V2: eight-step text-to-video with synchronized audio from 42 of the 50 MiniMax H3 transformer blocks. We removed the eight blocks whose removal changed video and audio predictions the least, replaced each block's AdaLN timestep projection with a shared rank-16 basis, and trained the result with eight-step DMD2 (checkpoint 300). Removing blocks makes the model faster and smaller but costs some quality; use FastH3 V2 when quality matters most.

  • Sampling: 8 DMD steps (999, 874, 749, 624, 500, 375, 250, 125), sparse attention keeping 20% of tiles. fastvideo_inference.json holds the schedule, which FastVideo reads automatically.
  • Text encoder: NVFP4 Qwen3-VL trimmed to the 50 layers H3 reads.
  • VAE: LynnReal lightweight video VAE with Kijai's INT8 weights; H3 audio VAE.

Blog post: FastH3 on Consumer Hardware 路 Code: FastVideo

  • Transformer: 11.1 GiB. Attention, MLP and sparse-attention gate linears in NVFP4, each with a static activation scale calibrated on 1,000 prompts across all eight steps. For the RTX 5090, RTX PRO 6000 and DGX Spark.

Full-quality counterpart in the same format: FastVideo/FastVideo-FastH3-8-Step-V2-NVFP4-Consumer.

Downloads last month
25
Safetensors
Model size
0.9B params
Tensor type
F32
路
BF16
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for FastVideo/FastVideo-FastH3-Trim-8-Step-NVFP4

Finetuned
(7)
this model