FastH3 Trim, 8-step, FP8, for the RTX 4090, RTX 3090 and smaller GPUs

FastH3 Trim is an experimental, smaller version of FastH3 V2: eight-step text-to-video with synchronized audio from 42 of the 50 MiniMax H3 transformer blocks. We removed the eight blocks whose removal changed video and audio predictions the least, replaced each block's AdaLN timestep projection with a shared rank-16 basis, and trained the result with eight-step DMD2 (checkpoint 300). Removing blocks makes the model faster and smaller but costs some quality; use FastH3 V2 when quality matters most.

  • Sampling: 8 DMD steps (999, 874, 749, 624, 500, 375, 250, 125), sparse attention keeping 20% of tiles. fastvideo_inference.json holds the schedule, which FastVideo reads automatically.
  • Text encoder: NVFP4 Qwen3-VL trimmed to the 50 layers H3 reads.
  • VAE: LynnReal lightweight video VAE with Kijai's INT8 weights; H3 audio VAE.

Blog post: FastH3 on Consumer Hardware 路 Code: FastVideo

  • Transformer: 19.9 GiB. Attention and MLP linears in FP8 E4M3 with one scale per output channel; activations quantized per token at run time. Runs on the RTX 4090 and, with layerwise offload, within 16 GB and 12 GB.

Full-quality counterpart in the same format: FastVideo/FastVideo-FastH3-8-Step-V2-FP8.

Downloads last month
55
Safetensors
Model size
19B params
Tensor type
F32
路
BF16
路
F8_E4M3
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for FastVideo/FastVideo-FastH3-Trim-8-Step-FP8

Finetuned
(7)
this model