Instructions to use Giniiki/FastWan2.2-TI2V-5B-mlx-q8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Giniiki/FastWan2.2-TI2V-5B-mlx-q8 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir FastWan2.2-TI2V-5B-mlx-q8 Giniiki/FastWan2.2-TI2V-5B-mlx-q8
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
FastWan 2.2 TI2V-5B โ MLX q8 (Pithos)
The image-to-video checkpoint Pithos downloads. DMD-distilled Wan 2.2
TI2V-5B (3 denoise steps, sigmas [1.0, 0.757, 0.522, 0.0], renoise,
guide 1), quantized for Apple Silicon.
| file | size | what |
|---|---|---|
model.safetensors |
5.4 GB | DiT, q8 (group 64) โ attention + FFN Linears quantized, per mlx-video's predicate |
t5_encoder_q8.safetensors |
6.4 GB | umT5-XXL, q8 (group 64) everywhere, scales kept float32 |
vae.safetensors |
2.8 GB | Wan 2.2 VAE (encoder + decoder), unquantized |
tokenizer.json + configs |
17 MB | umT5 tokenizer (T5Tokenizer / Unigram) |
config.json |
โ | model config incl. the fastwan_dmd recipe and quantization block |
Provenance: DiT quantized 2026-08-30 from
lBroth/FastWan2.2-TI2V-5B-MLX
(bf16, rev de2998ac446581895089ae363f4dcc0c02ca9ea2), itself converted from
FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers.
The T5 and VAE come from that same conversion (T5 quantized here); the
tokenizer files are copied from
Wan-AI/Wan2.2-TI2V-5B
google/umt5-xxl/. Quantization used mlx.nn.quantize via mlx-video's
converter (mlx 0.32.2).
Two facts written into Pithos's tests, kept here so a re-conversion does not lose them: the T5's quantization scales must stay float32 (bf16 scales measurably degrade the encoder, 0.023 โ 0.042 relative), and the DMD distill must be sampled at its trained 121-frame profile with renoise โ plain euler over its sigmas, or off-profile frame counts, dissolve the clip's tail.
License: Apache-2.0, inherited from the base model.
- Downloads last month
- 45
8-bit