Instructions to use vanch007/AuK-Flash-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use vanch007/AuK-Flash-MLX with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir AuK-Flash-MLX vanch007/AuK-Flash-MLX
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
π AuK-Flash-MLX (Native Apple Silicon Port)
π GitHub Source & Web Demo: vanch007/mlx-AuK Β· 8-Bit Quantized MLX Weights
This repository provides the native Apple Silicon MLX port of AuK-Flash, Tencent Hunyuan's 1.5B speech foundation model for ultra-fast generation and zero-shot speech/lyric editing on Apple Silicon.
- Main GitHub Repository: https://github.com/vanch007/mlx-AuK
- Original Project: Tencent-Hunyuan/AuK
- Architecture: 10 Double-Stream MMDiT Blocks + 20 Single-Stream DiT Blocks + Causal BigVGAN-Flow-VAE + Qwen2.5-Omni Thinker
- Inference Mode: 4-Step Distilled Flow Matching (DMD) with fixed CFG=0.0
- Sampling Rate: 24kHz Mono High-Fidelity Audio
Performance on Apple Silicon
| Metric | PyTorch MPS | Native MLX (This Model) | MLX 8-Bit Quantized |
|---|---|---|---|
| 4-Step DiT Latency (10s Audio) | 3.820s | 0.992s (3.85x faster) | 1.022s |
| Real-Time Factor (RTF) | 0.3820 | 0.0992 (10.08x real-time) | 0.1022 (9.79x real-time) |
| Backbone Memory | 5.70 GB | 5.70 GB | 0.56 GB (90.1% saving) |
Files Included
- : Full-precision MLX weights for the Flux2Edit diffusion transformer backbone (5.7GB).
- : Full-precision MLX weights for the BigVGAN Flow VAE (608MB).
- : Architecture and model hyperparameters.
Quickstart & Interactive Demo
For complete usage instructions, Python API, and the interactive A/B comparison web demo, visit: π https://github.com/vanch007/mlx-AuK
- Downloads last month
- 19
Hardware compatibility
Log In to add your hardware
Quantized