Hojo-TTS-Light-40M-INT8

int8 quantization of the decoder LM graph (Hojo-TTS-Light-40M-llm.onnx) from HojoAI/Hojo-TTS-Light-40M at revision c3cf21c77dc04b220aa4dabe20fc59e66a2a5485.

This repository intentionally holds only that one graph. The install pipeline resolves each snapshot file independently, so the remaining components (decoder, fine_local, voice bank, tokenizer, config) continue to come from the upstream snapshot; only the LM is served from here.

File

file bytes sha256
Hojo-TTS-Light-40M-llm.onnx 32358615 e43fc9b10ccfecc7044eded41bed88c1e5f81fdd3e955a200e340246eeb0ee2c

The upstream file at the revision above is BF16 (63138277 bytes) and is promoted to FP32 (125813561 bytes) at install time. This file replaces that intermediate step: it is already int8, so the promotion pass finds no BF16 tensors and leaves the numerics untouched (verified bitwise-identical step outputs).

How it was produced

pip install onnxruntime
python tool/quantize_hojo_tts_llm.py <installed-model-dir>

which runs:

onnxruntime.quantization.quantize_dynamic(
    source, target,
    weight_type=QuantType.QInt8,
    op_types_to_quantize=["MatMul", "Gemm"],
    per_channel=True,
    reduce_range=False,
)

Only the autoregressive decoder LM is quantized. The vocoder graphs (decoder, fine_local) are deliberately left at their original precision: dynamic weight-only quantization buys little on their convolutions, and they are where audible artifacts would come from.

Why this graph

The LM runs once per generated audio token โ€” roughly 150-190 session runs for a 35 character sentence โ€” so it dominates synthesis time. One decode step, measured on an i5-12600KF (16 logical cores, 6P+4E):

4 threads 8 threads size
fp32 21.2 ms 22.7 ms 125.8 MB
int8 8.9 ms 7.3 ms 32.4 MB

End to end through the synthesis pipeline the real-time factor for the same sentences falls from 1.20-1.32 to 0.57-0.67 โ€” from slower than real time to faster than real time, which is what lets a read-aloud queue keep up with speech.

Note that quantizing shifts the optimal thread count: 8 threads beat 4 for int8, while the reverse holds for fp32.

License

Apache-2.0, inherited from the base model (see LICENSE). This repository is a precision-only derivative of the upstream weights; the architecture is unchanged.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for nocmt/Hojo-TTS-Light-40M-INT8

Quantized
(1)
this model