VoxCPM2-MLX: Native Apple Silicon Model Weights

This repository provides 100% native Apple Silicon (MLX) ready weights for VoxCPM2, OpenBMB's 2B parameter foundation Text-to-Speech and Voice Cloning model supporting 30 languages and 48kHz studio quality output.

🌍 Supported Languages (30)

Turkish, English, Chinese, German, French, Spanish, Italian, Japanese, Korean, Arabic, Russian, Portuguese, Dutch, Polish, Swedish, Danish, Finnish, Norwegian, Greek, Hebrew, Hindi, Indonesian, Vietnamese, Thai, Tagalog, Swahili, Malay, Burmese, Khmer, Lao.

πŸ“¦ What's Included

  • model.safetensors: 2B Base & Residual Acoustic LM + Unified CFM (LocDiT)
  • audiovae.safetensors: Pre-converted, fused-weight AudioVAE V2 for 48kHz studio audio synthesis directly on Apple Silicon Metal GPU (no PyTorch required).
  • Tokenizers and model configurations.

πŸš€ Quick Start

Install the MLX package:

git clone https://github.com/hbasrisahin/VoxCPM2.git
cd VoxCPM2
pip install -e .

Run in Python:

from voxcpm2 import VoxCPM2Pipeline

# Automatically downloads and loads this MLX repository
pipe = VoxCPM2Pipeline.from_pretrained('hbasrisahin/VoxCPM2-MLX')

audio = pipe.generate(
    text='Merhaba! Bu ses Apple Silicon ΓΌzerinde saf MLX motoru ile ΓΌretildi.',
    inference_timesteps=10,
    seed=42,
)

import soundfile as sf
sf.write('output.wav', audio, pipe.sample_rate)

πŸ“„ License & Attribution

Based on the VoxCPM2 foundation model by OpenBMB, licensed under Apache-2.0.

Downloads last month
37
Safetensors
Model size
2B params
Tensor type
BF16
Β·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support