Audio8-TTS-Preview-0.1b-bf16 (MLX)

MLX conversion of Audio8/Audio8-TTS-Preview-0.1b (revision 1b17c91) for Apple Silicon — a 0.1B-parameter multilingual text-to-speech model with zero-shot voice cloning and a bundled 44.1 kHz neural codec. And then quantized to Q6 or Q4 for variable layers. Then finally merge codec into the model to get a less than 1GB model file.

  • Language model: bf16 (as published upstream)
  • Codec: fp32, converted from codec.pth to safetensors with weight-norm folded and conv weights in MLX (channels-last) layout
  • Architecture: arktts — a DualAR transformer (24-layer slow AR predicting one semantic token per frame, 4-layer fast AR predicting 10 codec codebooks per frame), inspired by Fish Audio S2 Pro

Conversion parity vs the PyTorch reference (fp32, CPU): unit/block outputs within 1e-4, reference-audio codec encoding 100% code-exact, greedy generation 100% token-exact over the validation utterance, decoded waveform max-abs 7.5e-6.

Usage

Currently can only be used in https://github.com/jason-ni/tf-assistant-app

Supported languages

Cantonese, Chinese, Dutch, English, French, German, Italian, Japanese, Korean, Polish, Spanish (per the upstream preview release).

License

Apache-2.0, following the upstream model. The bundled codec weights are part of the upstream repository and carry the same license.

Downloads last month
46
Safetensors
Model size
0.2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for plaincompute/audio8-tts-preview-0.1b-mx

Quantized
(2)
this model