NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Paper โข 2403.03100 โข Published โข 37
Converted to the Vokra GGUF format for Vokra, a zero-dependency speech-AI inference runtime.
This is a conversion, not a new model. The weights are the upstream ones; Vokra re-packages them so its runtime can memory-map them directly. Credit for the model belongs upstream โ see Source below.
| File | Size | SHA-256 |
|---|---|---|
naturalspeech3-facodec-v2.gguf |
428.4 MB | ee1d1e23266d6d2a898152d18bde156a2de008b8fe1eae9eeb392feca24c3084 |
# Download (any HTTP client works โ the file is a plain GGUF)
curl -L -o naturalspeech3-facodec-v2.gguf \
https://huggingface.co/vokra/naturalspeech3-facodec-v2/resolve/main/naturalspeech3-facodec-v2.gguf
vokra-cli run --model naturalspeech3-facodec-v2.gguf --input input.wav
| Field | Value |
|---|---|
| Architecture | facodec |
| Tensors | 806 |
| Upstream source | amphion/naturalspeech3_facodec v2 (Amphion NaturalSpeech 3 FACodec โ encoder_v2 + decoder_v2, factorized VQ codec, prosody 1cb + content 2cb + detail 3cb, 16 kHz, hop 200 โ 80 tok/s, arXiv:2403.03100, apache-2.0) |
| Upstream licence | apache-2.0 |
| Licence class | permissive |
| Registry model id | naturalspeech3-facodec-v2 |
| Vokra GGUF schema | 1 |
| Converted by | vokra-core 0.1.0-alpha.0 |
Every row above is read out of this file's own vokra.* metadata, so the card cannot claim something the artifact does not carry.
The weights are distributed under apache-2.0, unchanged from upstream. Conversion does not alter the licence, and your obligations run to the upstream author.
shasum -a 256 naturalspeech3-facodec-v2.gguf
# expect: ee1d1e23266d6d2a898152d18bde156a2de008b8fe1eae9eeb392feca24c3084
We're not able to determine the quantization variants.