EditVoice โ Official Model Weights
This is the official model weights repository for EditVoice: Variable-Length Non-Autoregressive Zero-Shot TTS and Speech Editing with Edit Flows. It provides the inference checkpoints and Edit Flow initial distribution for zero-shot TTS and speech editing in noisy environments. The official code repository contains the inference implementation, and audio examples are available in the EditVoice demo.
Files and Provenance
| File | Purpose | Source and license |
|---|---|---|
checkpoints/editvoice_lm.pt |
Edit Flow semantic model for TTS and speech editing | EditVoice; Apache-2.0 |
checkpoints/kimi-flow.pt |
Token-to-mel Flow Matching model fine-tuned for speech editing in noisy environments; paired with checkpoints/kimi-bigvgan/ |
EditVoice fine-tuning from CosyVoice2 flow.pt; Apache-2.0 |
data/empirical.pkl |
6,561-class initial distribution $p_0$ used by Edit Flows | EditVoice training-data statistics; Apache-2.0 |
checkpoints/flow.pt |
Token-to-mel Flow Matching model for zero-shot TTS; paired with checkpoints/hift.pt |
CosyVoice2-0.5B; Apache-2.0 |
checkpoints/hift.pt |
Vocoder for zero-shot TTS; paired with checkpoints/flow.pt |
CosyVoice2-0.5B; Apache-2.0 |
checkpoints/speech_tokenizer_v2.onnx |
Speech tokenizer | CosyVoice2-0.5B; Apache-2.0 |
checkpoints/campplus.onnx |
Speaker embedding extractor | CosyVoice2-0.5B; Apache-2.0 |
checkpoints/kimi-bigvgan/config.json, checkpoints/kimi-bigvgan/model.pt |
Vocoder for speech editing in noisy environments; paired with checkpoints/kimi-flow.pt |
vocoder/ from Kimi-Audio-7B-Instruct; MIT |
The ttsfrd wheel and its English resources are not included; obtain them separately from an authorized source.
Download and Use
With the Hugging Face CLI installed, download the files and verify their checksums:
hf download DOOD02/EditVoice --local-dir model_assets
(cd model_assets && sha256sum -c SHA256SUMS)
For editvoice-tts, pass model_assets/checkpoints/editvoice_lm.pt, flow.pt, hift.pt, speech_tokenizer_v2.onnx, campplus.onnx, and model_assets/data/empirical.pkl to their corresponding checkpoint and resource arguments. For editvoice-edit, use the same semantic model, tokenizer, speaker model, and distribution, together with model_assets/checkpoints/kimi-flow.pt and model_assets/checkpoints/kimi-bigvgan/. See the EditVoice code repository README for complete commands and input formats.
License
This repository contains files under Apache-2.0 and MIT; no single license applies to every file. See LICENSE for the file-by-file license mapping and licenses/ for the license texts. Third-party weights retain their original authorship and licenses.