FastH3 V2, 8-step, NVFP4 for one GPU

FastH3 V2 (8 distilled steps, video with synchronized audio) packaged for a single Blackwell GPU (RTX 5090, RTX PRO 6000) or a DGX Spark.

  • Transformer: FastH3 V2 (50 blocks). MLP linears keep the calibrated NVFP4 weights and activation scales from FastVideo/FastVideo-FastH3-8-Step-V2-NVFP4; attention projections and sparse-attention gates are also stored in NVFP4.
  • Text encoder: NVFP4 Qwen3-VL trimmed to the 50 layers H3 reads.
  • VAE: LynnReal lightweight video VAE with Kijai's INT8 weights; H3 audio VAE.
  • Sampling: fastvideo_inference.json holds the 8-step DMD schedule, which FastVideo reads automatically.

Machines with less than 64 GB of system RAM

The 50 AdaLN timestep projections are 26 GB in BF16. transformer/adaln_tables.pt (155 MB) holds their precomputed outputs for the 8-step schedule, so FastVideo can skip loading them:

import os
from huggingface_hub import hf_hub_download

os.environ["FASTVIDEO_H3_ADALN_TABLE"] = hf_hub_download(
    "FastVideo/FastVideo-FastH3-8-Step-V2-NVFP4-Consumer", "transformer/adaln_tables.pt")

The table covers only the shipped schedule (fastvideo_inference.json); FastVideo raises an error for other timesteps.

For multi-GPU data-center serving, use FastVideo/FastVideo-FastH3-8-Step-V2-NVFP4. Smaller, faster experimental variant: FastVideo/FastVideo-FastH3-Trim-8-Step-NVFP4.

Downloads last month
41
Safetensors
Model size
14B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FastVideo/FastVideo-FastH3-8-Step-V2-NVFP4-Consumer

Finetuned
(7)
this model