Qwen3-TTS-12Hz-0.6B-CustomVoice for Intel Core Ultra
Maintained by AtomGradient / 质子梯度(北京)科技有限公司.
Model artifacts for local speech synthesis on Intel hardware. The evaluated configuration uses the Qwen3-TTS 0.6B CustomVoice model, with INT8 weight storage and floating-point components. INT8 does not describe every computation in the system.
Model and device specifications
| Item | Specification |
|---|---|
| Base model | Qwen3-TTS-12Hz-0.6B-CustomVoice |
| Model size class | 0.6B |
| Output audio | 24 kHz, mono, 16-bit PCM in the evaluated service |
| Voices | 9 preset voices; default evaluated voice: Vivian |
| Validated device | Intel Core Ultra 5 225H, Intel Arc integrated graphics, 32 GiB DDR5 |
| Evaluation focus | Chinese dialogue and storytelling |
Voices: Vivian, Serena, Dylan, Eric, Aiden, Ryan, Ono_anna, Sohee and Uncle_fu. The language metadata describes the base model's language coverage; it is not a claim of equivalent validation across those languages.
Observed performance
The following results were measured on the device above using the AtomGradient speech runtime. They describe the integrated service, not a performance guarantee for the model files alone.
| Measurement | Observed result | Test scope |
|---|---|---|
| First audio block sent | Median 0.369 s; range 0.347–0.497 s | 10 dialogue sessions, 2026-09-12 |
| Real-time factor (RTF) | Median 0.716; range 0.703–0.851 | 18 completed synthesis segments in the same window |
| Dialogue completion | 8 completed sessions, 2 intentional cancellations; subsequent sessions completed | Same field-test window |
| Long response | 367 characters, 87.84 seconds of generated audio | One completed Chinese response |
| Continuous operation | 30 minutes, 357 utterances, 9 voices, no failed utterances | Separate TTS-only terminal test, 2026-09-08 |
RTF is generation time divided by generated audio duration; below 1.0 means generation is faster than playback. First-block timing starts when the local service receives text and ends when it sends the first audio block. It excludes LLM waiting and acoustic playback latency. Streaming output does not require the entire response to finish generating before audio becomes available.
Scope and limitations
- These are measurements on one device and specific test inputs, not population statistics or a service-level guarantee.
- The field test and the earlier continuous-operation test are separate observations and must not be combined into one benchmark.
- No independent naturalness score or lossless-quantization claim is made. English synthesis and additional language or voice scenarios need further quality evaluation.
- This repository contains model artifacts and associated configuration. The AtomGradient runtime, installation package and product integration are delivered separately; downloading this repository alone does not reproduce the complete demonstrated service.
For deployment and integration, contact AtomGradient.
License and attribution
The base model is Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice. Model artifacts are distributed under Apache-2.0, following the base model's license. AtomGradient's copyright covers its own contributions and does not extend to the upstream Qwen model or third-party components. See THIRD_PARTY.md for attribution.
Model tree for AtomGradient/Qwen3-TTS-12Hz-0.6B-CustomVoice-streaming-int8-ov
Base model
Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice