πŸš€ AuK-Flash-MLX (Native Apple Silicon Port)

πŸ”— GitHub Source & Web Demo: vanch007/mlx-AuK Β· 8-Bit Quantized MLX Weights

This repository provides the native Apple Silicon MLX port of AuK-Flash, Tencent Hunyuan's 1.5B speech foundation model for ultra-fast generation and zero-shot speech/lyric editing on Apple Silicon.

  • Main GitHub Repository: https://github.com/vanch007/mlx-AuK
  • Original Project: Tencent-Hunyuan/AuK
  • Architecture: 10 Double-Stream MMDiT Blocks + 20 Single-Stream DiT Blocks + Causal BigVGAN-Flow-VAE + Qwen2.5-Omni Thinker
  • Inference Mode: 4-Step Distilled Flow Matching (DMD) with fixed CFG=0.0
  • Sampling Rate: 24kHz Mono High-Fidelity Audio

Performance on Apple Silicon

Metric PyTorch MPS Native MLX (This Model) MLX 8-Bit Quantized
4-Step DiT Latency (10s Audio) 3.820s 0.992s (3.85x faster) 1.022s
Real-Time Factor (RTF) 0.3820 0.0992 (10.08x real-time) 0.1022 (9.79x real-time)
Backbone Memory 5.70 GB 5.70 GB 0.56 GB (90.1% saving)

Files Included

  • : Full-precision MLX weights for the Flux2Edit diffusion transformer backbone (5.7GB).
  • : Full-precision MLX weights for the BigVGAN Flow VAE (608MB).
  • : Architecture and model hyperparameters.

Quickstart & Interactive Demo

For complete usage instructions, Python API, and the interactive A/B comparison web demo, visit: πŸ‘‰ https://github.com/vanch007/mlx-AuK

Downloads last month
19
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support