DocXAssist voices (Kokoro, ONNX)

Text-to-speech voices used by the DocXAssist desktop app (https://docxassist.com) for Spanish and German. Both are Kokoro-82M (StyleTTS 2 architecture) models in ONNX format. Large weights are stored as float16 and are expanded to float32 by ONNX Runtime when the model is loaded (same speed and quality as the float32 model, half the file size).

Folder Source License
es/ hexgrad/Kokoro-82M v1.0 via onnx-community/Kokoro-82M-v1.0-ONNX; voices ef_dora, em_alex Apache-2.0
de/ Thorsten-Voice/Kokoro (German fine-tune of Kokoro-82M on the CC0 Thorsten-Voice dataset, recipe by kikiri-tts); voice thorsten Apache-2.0

Changes made by DocXAssist

  • de/model.onnx: exported from the PyTorch checkpoint model.pth to ONNX (opset 17, disable_complex=True).
  • Both models: weights with ≥ 1024 values converted to float16 with a Cast node back to float32.
  • Voice style tensors saved as raw float32 arrays (*.bin, 510 × 1 × 256). vocab.json is the phoneme vocabulary.

Inputs

input_ids (int64, [1, n], phoneme ids wrapped in 0 … 0), style (float32, [1, 256], row n − 3 of the voice tensor), speed (float32, [1]). Output: waveform (float32, 24 kHz). Phonemes come from eSpeak NG in IPA with ties, mapped as in misaki EspeakG2P.

All credit for the models goes to their authors: hexgrad (Kokoro), Thorsten Müller (Thorsten-Voice), semidark (kikiri-tts).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support