FastH3 V2, 8-step, FP8

FastH3 V2 (8 distilled steps, video with synchronized audio) with an FP8 transformer for GPUs without FP4 tensor cores, such as the RTX 4090 and RTX 3090, and for 16 GB and 12 GB memory tiers.

  • Transformer: FastH3 V2 (50 blocks). Attention and MLP linears in FP8 E4M3 with one scale per output channel; activations are quantized per token at run time. Everything else stays BF16.
  • Text encoder: NVFP4 Qwen3-VL trimmed to the 50 layers H3 reads (dequantized per layer on GPUs without FP4).
  • VAE: LynnReal lightweight video VAE with Kijai's INT8 weights; H3 audio VAE.
  • Sampling: fastvideo_inference.json holds the 8-step DMD schedule, which FastVideo reads automatically.

Source weights: FastVideo/FastVideo-FastH3-8-Step-V2. Smaller, faster experimental variant: FastVideo/FastVideo-FastH3-Trim-8-Step-FP8.

Downloads last month
-
Safetensors
Model size
35B params
Tensor type
F32
路
BF16
路
F8_E4M3
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for FastVideo/FastVideo-FastH3-8-Step-V2-FP8

Finetuned
(7)
this model