VoTSpeech Model Weights

This repository contains the VoTSpeech model weights and the configuration and tokenizer assets required to load them. The checkpoint is based on dots-studio/dots.tts-soar and generates 48 kHz speech from text and natural-language voice instructions.

Code and usage

For inference code, installation, and usage instructions, see ywbn/VoTSpeech.

Weight files

  • model.safetensors: VoTSpeech model weights
  • vocoder.safetensors: 48 kHz vocoder weights
  • speaker_encoder.safetensors: speaker encoder weights
  • latent_stats.pt: voice-latent normalization statistics
  • config.json, llm_config.json: model configuration
  • tokenizer.json, tokenizer_config.json, chat_template.jinja: tokenizer assets
  • SHA256SUMS: release integrity checksums

License and copyright

The VoTSpeech model weights are licensed under CC BY-NC 4.0, allowing non-commercial use, sharing, and adaptation with attribution and an indication of any changes. When sharing the weights or adaptations, credit VoTSpeech, link to this repository and the license, and identify your modifications. The official legal text sets out the full terms.

This release provides model weights and supporting loading assets, not the training datasets or original recordings. Rights in third-party recordings, texts, performances, and other source materials remain with their respective rights holders. The weight license does not grant permission to redistribute those materials or waive any applicable privacy, publicity, or personality rights. It does not replace any separately applicable data-use agreements.

Intended use and responsible use

VoTSpeech is intended to support research and evaluation of instruction-guided voice design and expressive speech synthesis. Commercial use of the licensed weights is not permitted under CC BY-NC 4.0.

As responsible-use guidance, we ask users to clearly identify publicly shared outputs as synthetic speech, obtain any permissions needed when imitating an identifiable voice, and avoid deceptive impersonation, fraud, harassment, or privacy-invasive applications. These recommendations do not add restrictions to the CC BY-NC 4.0 license; applicable laws and third-party rights still apply.

The model may produce inaccurate pronunciations, unintended voice attributes, or audio artifacts. It is provided as is, without warranties, to the extent permitted by law. Users are responsible for assessing outputs and ensuring that their use is lawful and appropriate.

For licensing or rights concerns, please contact the maintainers through the Hugging Face community page. Please do not post sensitive personal information publicly.

Downloads last month
-
Safetensors
Model size
2B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ywbn16/VoTSpeech

Finetuned
(3)
this model