YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Qwen3-TTS benchmarks on RTX PRO 6000 Blackwell

Two studies live in this repo — note the model size for every table:

study model date report raw data
TRT vs native vocoder A/B (FastAPI + Silero path) 1.7B Qwen3-TTS-RU-CallCenter-stressed 2026-07-14 this README (below) + VOCODER_RUNTIME_BENCH.md metrics/native-eager-fastapi-silero-c10-c60/, metrics/trt-fastapi-silero-c10-c60/
Load test + 1.7B comparison + TRT-vs-eager A/B (direct vLLM path) 0.6B Qwen/Qwen3-TTS-12Hz-0.6B-Base 2026-07-16 QWEN06B_BENCH.md metrics/qwen06b-trt-c1-c50/, metrics/qwen06b-trt-high-c30-c64/, metrics/qwen06b-eager-c1-c50/

[1.7B] TensorRT Vocoder vs Native PyTorch Vocoder with FastAPI and Silero Stress

Exact A/B measurement date: 2026-07-14. Hardware: NVIDIA RTX PRO 6000 Blackwell Server Edition. Model: Qwen3-TTS-RU-CallCenter-stressed, Base voice clone with 10_ekaterina.wav, PCM16 mono at 24 kHz.

Result

The TensorRT vocoder sustained every tested level from 10 through 60 concurrent sessions within SLO. The native PyTorch/eager vocoder passed through concurrency 30 and failed at 40, 50, and 60.

Both runs used the complete public service path, including Russian text normalization and automatic stress placement:

DVC replay -> FastAPI :8000 -> normalization -> Silero Stress -> vLLM -> vocoder

The only intentional inference-runtime variable was:

  • TRT run: OMNI_TRT_VOCODER=1, code2wav decoded by TensorRT in Triton.
  • Native run: OMNI_TRT_VOCODER=0, code2wav decoded by the original in-process PyTorch/eager implementation.

Exact workload controls

Both runs used:

  • the same 100 private-HF call scenarios and seed 42;
  • the same 10_ekaterina.wav and transcript reference;
  • the same qwen3_tts_bigmem.yaml deployment configuration;
  • the same sentence chunker and LLM simulator;
  • two in-flight TTS chunks per session;
  • real recorded user-pause pace (user_time_scale: 1.0);
  • 90-second sustained windows at concurrency 10, 20, 30, 40, 50, and 60;
  • the same SLO thresholds and 200 ms audible-gap boundary.

TRT-path verification: the Triton code2wav inference counter increased during the TRT run. Native-path verification: it remained exactly 30674 -> 30674 through the complete native run. After the experiment, the TRT service was restored and the counter increased again on a control request.

Gap definitions

  • Micro-gap: a playback-buffer underrun shorter than 200 ms. It is measured, written into the WAV timeline, and logged, but does not count as an audible SLO stutter.
  • Audible gap: a playback-buffer underrun of at least 200 ms.
  • Total gaps: micro-gaps plus audible gaps.
  • Gap frequency: gap events per 100 completed turns. Turns can contain more than one gap, so this is an event rate rather than a percentage of turns.
  • Audible-gap turns: percentage of turns containing at least one gap of 200 ms or longer.
  • Initial silence before the first PCM bytes is represented by TTFA/E2E and is not counted as a mid-phrase gap.

Primary comparison: gap scaling

Concurrency TRT turns TRT total gaps TRT gaps / 100 turns TRT audible gaps Native turns Native total gaps Native gaps / 100 turns Native audible gaps Native audible-gap turns
10 60 0 0.00 0 61 0 0.00 0 0.00%
20 119 0 0.00 0 119 2 1.68 0 0.00%
30 185 0 0.00 0 178 11 6.18 6 3.37%
40 246 2 0.81 0 246 36 14.63 14 5.69%
50 310 7 2.26 0 305 72 23.61 28 8.52%
60 374 18 4.81 0 358 157 43.85 54 12.57%

At concurrency 60, the native vocoder produced one playback gap per 2.28 turns on average. The TRT vocoder produced one short micro-gap per 20.78 turns and no audible gaps.

Gap duration detail

Concurrency TRT micro-gaps TRT micro p95, ms TRT max gap, ms Native micro-gaps Native micro p95, ms Native audible gaps Native audible p95, ms Native max gap, ms
10 0 - - 0 - 0 - -
20 0 - - 2 195.6 0 - 195.6
30 0 - - 5 188.9 6 360.2 360.2
40 2 33.8 33.8 22 166.0 14 910.8 911.6
50 7 53.8 53.8 44 193.2 28 1,217.8 1,846.1
60 18 71.0 110.3 103 172.4 54 1,235.6 1,242.0

The TRT run's worst gap across all levels was 110.3 ms. The native run's worst gap was 1,846.1 ms at concurrency 50.

Latency and throughput

Concurrency TRT TTFA p95, ms Native TTFA p95, ms TRT E2E p95, ms Native E2E p95, ms TRT audio x Native audio x TRT SLO Native SLO
10 173.5 202.5 1,120.7 1,126.1 7.64 7.79 PASS PASS
20 259.9 440.8 1,210.3 1,280.7 16.87 16.21 PASS PASS
30 278.6 489.6 1,179.3 1,242.6 23.51 23.82 PASS PASS
40 504.6 1,151.0 1,243.8 2,004.0 31.13 31.01 PASS FAIL
50 522.4 1,720.7 1,258.9 2,315.6 39.05 39.02 PASS FAIL
60 578.2 2,816.9 1,319.6 3,658.6 47.40 45.16 PASS FAIL

Throughput remains similar through concurrency 50, but the native vocoder's queueing latency and playback discontinuities grow sharply. At concurrency 60, native TTFA p95 is 4.87 times the TRT value.

Detailed run tables

TensorRT vocoder

Concurrency Turns Micro-gaps Audible gaps Gap events / 100 turns TTFA p50/p95/p99, ms E2E p95, ms Errors SLO
10 60 0 0 0.00 130.2 / 173.5 / 194.9 1,120.7 0 PASS
20 119 0 0 0.00 148.0 / 259.9 / 299.0 1,210.3 0 PASS
30 185 0 0 0.00 153.9 / 278.6 / 331.1 1,179.3 0 PASS
40 246 2 0 0.81 167.1 / 504.6 / 642.4 1,243.8 0 PASS
50 310 7 0 2.26 176.1 / 522.4 / 694.0 1,258.9 0 PASS
60 374 18 0 4.81 197.6 / 578.2 / 845.8 1,319.6 0 PASS

Measured TRT capacity lower bound: at least 60 concurrent sessions. A level above 60 was not included in this exact A/B run.

Native PyTorch/eager vocoder

Concurrency Turns Micro-gaps Audible gaps Gap events / 100 turns TTFA p50/p95/p99, ms E2E p95, ms Errors SLO
10 61 0 0 0.00 143.9 / 202.5 / 253.9 1,126.1 0 PASS
20 119 2 0 1.68 181.3 / 440.8 / 537.0 1,280.7 0 PASS
30 178 5 6 6.18 202.6 / 489.6 / 706.6 1,242.6 0 PASS
40 246 22 14 14.63 235.7 / 1,151.0 / 3,334.2 2,004.0 0 FAIL
50 305 44 28 23.61 322.3 / 1,720.7 / 3,542.4 2,315.6 0 FAIL
60 358 103 54 43.85 460.9 / 2,816.9 / 5,477.3 3,658.6 0 FAIL

Measured native capacity under the configured SLO: 30 concurrent sessions.

Conclusions

  1. Silero Stress is compatible with both vocoder implementations: neither run produced HTTP or TTS errors, and both returned valid streaming PCM.
  2. The original native vocoder remains acceptable through concurrency 30, but playback gaps begin at 20 and audible gaps begin at 30.
  3. Native latency crosses the SLO at concurrency 40 and degrades rapidly at 50/60 despite similar aggregate audio throughput.
  4. TensorRT eliminates all audible gaps in the tested 10-to-60 range and keeps TTFA/E2E p95 within SLO at every level.
  5. The measured SLO capacity improvement is 30 -> at least 60 concurrent sessions.

Data provenance

TRT aggregate artifacts:

  • experiments/fastapi-stress-c10-c60-90s-20260714/metrics.json
  • experiments/fastapi-stress-c10-c60-90s-20260714/env.json
  • experiments/fastapi-stress-c10-c60-90s-20260714/params_used.yaml
  • experiments/fastapi-stress-c10-c60-90s-20260714/gpu_dmon.log

Native aggregate artifacts:

  • experiments/fastapi-stress-native-eager-c10-c60-90s-20260714/metrics.json
  • experiments/fastapi-stress-native-eager-c10-c60-90s-20260714/env.json
  • experiments/fastapi-stress-native-eager-c10-c60-90s-20260714/params_used.yaml
  • experiments/fastapi-stress-native-eager-c10-c60-90s-20260714/gpu_dmon.log

Private-corpus artifacts (events.jsonl, turns.jsonl, WAV recordings, and the HTML report containing turn text) are intentionally kept local and are not published to the public Hugging Face repository.

  • Scenario SHA-256: 627261a0ada620c562449c53ffc6ee546eb29720d68ba3327c95a22d506b77cb
  • Reference WAV SHA-256: d913ffab4a16c5a017a36ad2d00a8352fcdb190396d64f37d5a5d91466dec673

SLO thresholds: TTS TTFA p95 <= 800 ms, E2E first-audio p95 <= 2,000 ms, audible underruns <= 0.2 per turn, and errors <= 1%.


[0.6B] Update 2026-07-16: Qwen3-TTS-12Hz-0.6B-Base study

Everything in this section was measured on the 0.6B model (Qwen/Qwen3-TTS-12Hz-0.6B-Base), not the 1.7B production finetune.

Full report: QWEN06B_BENCH.md. Highlights (same GPU, same rig, shipped qwen3_tts_gb10.yaml config):

  • 0.6B swap is weight-only (identical codec dims, engines auto-rebuild).
  • Capacity within SLO: >= 50 concurrent calls (same as 1.7B) at 25-40 % lower TTFA: p95 at c50 350 ms vs 594 ms (1.7B).
  • TRT-vs-eager A/B on 0.6B: SLO capacity 50 vs 30; at c50 eager TTFA p95 2344 ms vs 350 ms TRT (x6.7) - the TRT vocoder wins regardless of talker size (the vocoder is the same in both checkpoints).
  • Raw data: metrics/qwen06b-trt-c1-c50/, metrics/qwen06b-trt-high-c30-c64/, metrics/qwen06b-eager-c1-c50/.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support