πŸš€ AuK-Base-MLX (Native Apple Silicon Port)

πŸ”— GitHub Source & Web Demo: vanch007/mlx-AuK Β· 8-Bit Quantized Base Weights Β· 4-Step AuK-Flash

This repository provides the native Apple Silicon MLX port of the full AuK Base foundation model, Tencent Hunyuan's 1.5B speech foundation model designed for high-fidelity speech generation, voice cloning, and zero-shot speech editing with configurable NFE steps and CFG guidance.

  • Main GitHub Repository: https://github.com/vanch007/mlx-AuK
  • Original Project: Tencent-Hunyuan/AuK
  • Architecture: 10 Double-Stream MMDiT Blocks + 20 Single-Stream DiT Blocks + Causal BigVGAN-Flow-VAE + Qwen2.5-Omni Thinker
  • Inference Mode: Configurable NFE (default 32 steps) & CFG strength (default 2.0)
  • Sampling Rate: 24kHz Mono High-Fidelity Audio

Performance on Apple Silicon

Metric PyTorch MPS Native MLX (This Model) Native MLX 8-Bit
32-Step Sampling Latency 29.8s 7.92s 8.15s
Real-Time Factor (RTF) 2.98 0.792 (1.26x real-time) 0.815
Backbone Memory 5.70 GB 5.70 GB 0.56 GB

Files Included

  • : Full-precision MLX weights for the AuK Base Flux2Edit backbone (5.7GB).
  • : Full-precision MLX weights for BigVGAN Flow VAE (608MB).
  • : Architecture and sampling parameters.

Quickstart

Downloads last month
38
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support