IndexTTS-2.5 ONNX
This model is used by https://github.com/mercallureAI/local-multimodal-infra .
This repository provides an FP16 ONNX export of IndexTeam/IndexTTS-2.5.
Original model:
https://huggingface.co/IndexTeam/IndexTTS-2.5
Original project:
https://github.com/index-tts/index-tts
This is not an official IndexTTS release. It is intended for ONNX Runtime inference testing and deployment experiments.
Description
IndexTTS-2.5 is a zero-shot text-to-speech model that clones a voice from a single reference clip. It supports Chinese, English, Japanese, Spanish and Arabic, with emotion control and speaking-speed control.
The package was built from the original IndexTTS-2.5 PyTorch weights in two steps:
Export: DakeQQ's exporter, used as is. The IndexTTS2 exporter (
Index_TTS/v2) from DakeQQ/Text-to-Speech-TTS-ONNX was run without modification at commitf491c82. The graph split, the graph inputs and outputs, and the FP16 KV cache all come from this exporter. Our wrapper script made one change, and it is totransformersrather than to the exporter: it mapsdtype=totorch_dtype=so the exporter runs on thetransformers4.52 that the officialuv.lockpins.FP16 conversion and packaging: our own code. DakeQQ's
Optimize_ONNX.pyand its FP16 plan were not used. Our script:- converts the graphs with ONNX Runtime's
convert_float_to_float16, keeping graph inputs and outputs in float32; - fuses the attention in the flow-matching DiT into
com.microsoft.MultiHeadAttention; - zeroes a token-id sum in the Synthesis graph that overflows FP16;
- removes redundant casts;
- drops the graphs the runtime does not use.
Weights shared between graphs are then deduplicated into
IndexTTS_SharedInitializers.onnx.data(1,511 tensors, 2.78 GB) with DakeQQ'sShared_Weights.bundle_shared_initializers, used unmodified. Our script writestokenizer.jsonandmanifest.json.- converts the graphs with ONNX Runtime's
The graphs cover reference-audio processing, the GPT backbone, the flow-matching mel decoder and the BigVGAN vocoder. Text normalization, tokenization, segmentation of long text and the sampling loop run outside the graphs, in your own code.
Files
| File | Precision | Role |
|---|---|---|
IndexTTS2_ReferencePreprocess.onnx |
Mixed: w2v-BERT encoder FP16, audio front end and speaker/mel path FP32 | Reference audio (22.05 kHz mono) to w2v-BERT semantic features, CAMPPlus speaker style and the reference mel condition |
IndexTTS2_Conditioning.onnx |
FP16 | Speaker latent and emotion vector, from the speaker and emotion reference features, the 8-way emotion weights and emotion_alpha |
IndexTTS2_TargetPrefill_sampling.onnx |
FP16 | GPT prefill over the text tokens and language_id; returns the FP16 KV cache and the first sampled mel code |
IndexTTS2_DecodeStep_sampling.onnx |
FP16 | One autoregressive step with top-k, top-p, temperature and repetition penalty; repeat until the stop code 8193 |
IndexTTS2_Synthesis.onnx |
FP16 | Mel codes to the flow-matching condition; applies duration_factor and cfg_rate |
IndexTTS2_CFMEstimator.onnx |
FP16 | One flow-matching step; run 25 times (cfm_steps) |
IndexTTS2_Decoder.onnx |
FP16 | BigVGAN vocoder: mel to 22.05 kHz waveform |
IndexTTS2_Metadata.onnx |
Carries the package contract as ONNX metadata | |
IndexTTS_SharedInitializers.onnx + .onnx.data |
FP16, some FP32 | Weight store referenced by the graphs above |
tokenizer.json |
Tokenizer description: split pattern, special tokens and language ids | |
multilingual_zh_ja_yue_char_del.tiktoken |
Text vocabulary, copied from the original model | |
manifest.json |
Package contract, runtime constants, shared-tensor layout and SHA-256 of every file | |
LICENSE, LICENSE_ZH |
bilibili Model Use License Agreement (English and Chinese) | |
config.json |
Package summary; the Hub counts downloads by requests for this file. Not read at run time |
Keep all files in one directory: the graphs load their weights from IndexTTS_SharedInitializers.onnx.data by relative path.
Runtime notes
- Requires ONNX Runtime:
IndexTTS2_CFMEstimator.onnxuses thecom.microsoft.MultiHeadAttentioncontrib operator. The graphs use opset 20. - Graph inputs and outputs are float32, except the KV cache, which is float16.
- Disable the ONNX Runtime optimizers
CastFloat16TransformerandFuseFp16InitializerToFp32NodeTransformerwhen creating sessions (listed inmanifest.jsonunderdisabled_optimizers). - FP16 targets NVIDIA GPUs (CUDA execution provider); plan for about 6 GB of VRAM. CPU works but is slow.
- Audio input and output are 22,050 Hz. Text is at most 600 tokens per segment; split longer text and join the segments.
- Emotion is an 8-float vector in the order
[happy, angry, sad, afraid, disgusted, melancholic, surprised, calm], scaled byemotion_alpha. Emotion from a text description is not included: the QwenEmotion graphs were not exported. - Only the sampling variants of the prefill and decode graphs are included; greedy decoding is not.
- Pronunciation can be fixed inline as
<word|READING>, using Pinyin with tone numbers, CMU phonemes or Kana, as in the original project.
Source
| Item | Revision |
|---|---|
IndexTeam/IndexTTS-2.5 weights |
c39ce5ba981572cb187443877ff559dfb246ce63 |
index-tts/index-tts code |
ee40fa7d6c6b8a2c7f06105f9f1e65775b74868c |
DakeQQ/Text-to-Speech-TTS-ONNX exporter (unmodified) |
f491c82ee6b29dd760c3c00e25552be8925358a7 |
DakeQQ/Text-to-Speech-TTS-ONNX is licensed under Apache-2.0. Only its output is distributed here; this repository contains none of its code.
The graphs also contain weights of the auxiliary models that IndexTTS-2.5 loads at runtime: facebook/w2v-bert-2.0, funasr/campplus and nvidia/bigvgan_v2_22khz_80band_256x. Their own licenses also apply.
Testing
The package was run end to end with ONNX Runtime from Rust on Chinese, English, Japanese and Spanish text, and a Chinese sample was transcribed back with ASR as a smoke check. Arabic has not been tested. No listening test or benchmark against the PyTorch model has been run, and FP16 output is not bit-identical to the original.
License
This repository is derived from IndexTeam/IndexTTS-2.5 and is distributed under the original bilibili Model Use License Agreement. See LICENSE and LICENSE_ZH.
The license is not Apache-2.0. Among other terms, organizations above the user or revenue thresholds in section 2.2 must request a separate license from bilibili, a copy of the agreement must be kept with every copy of the model, and downstream recipients must comply with it.
All rights to the original model, code and training work belong to the original authors.
Disclaimer
This is an unofficial ONNX export of IndexTTS-2.5.
It is not affiliated with, maintained by, or endorsed by the original IndexTTS authors.
Do not use this model to clone a person's voice without their consent, or for fraud, impersonation or misinformation.
Citation
@misc{li2026indextts25technicalreport,
title={IndexTTS 2.5 Technical Report},
author={Yunpei Li and Xun Zhou and Jinchao Wang and Lu Wang and Yong Wu and Siyi Zhou and Yiquan Zhou and Yining Wang and Yaogen Yang and Zhetao Hu and Shiyao Duan and Jiacheng Xu and Bin Xia and Jingchen Shu},
year={2026},
eprint={2601.03888},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2601.03888},
}
- Downloads last month
- 15
Model tree for ModaLeap/indextts-2.5-onnx
Base model
IndexTeam/IndexTTS-2.5