DocXAssist voices (Kokoro, ONNX)
Text-to-speech voices used by the DocXAssist desktop app (https://docxassist.com) for Spanish and German. Both are Kokoro-82M (StyleTTS 2 architecture) models in ONNX format. Large weights are stored as float16 and are expanded to float32 by ONNX Runtime when the model is loaded (same speed and quality as the float32 model, half the file size).
| Folder | Source | License |
|---|---|---|
es/ |
hexgrad/Kokoro-82M v1.0 via onnx-community/Kokoro-82M-v1.0-ONNX; voices ef_dora, em_alex |
Apache-2.0 |
de/ |
Thorsten-Voice/Kokoro (German fine-tune of Kokoro-82M on the CC0 Thorsten-Voice dataset, recipe by kikiri-tts); voice thorsten |
Apache-2.0 |
Changes made by DocXAssist
de/model.onnx: exported from the PyTorch checkpointmodel.pthto ONNX (opset 17,disable_complex=True).- Both models: weights with ≥ 1024 values converted to float16 with a Cast node back to float32.
- Voice style tensors saved as raw float32 arrays (
*.bin, 510 × 1 × 256).vocab.jsonis the phoneme vocabulary.
Inputs
input_ids (int64, [1, n], phoneme ids wrapped in 0 … 0), style (float32, [1, 256], row n − 3 of the
voice tensor), speed (float32, [1]). Output: waveform (float32, 24 kHz). Phonemes come from eSpeak NG in
IPA with ties, mapped as in misaki EspeakG2P.
All credit for the models goes to their authors: hexgrad (Kokoro), Thorsten Müller (Thorsten-Voice), semidark (kikiri-tts).