Qwen3-ASR-0.6B for Intel Core Ultra
Maintained by AtomGradient / 质子梯度(北京)科技有限公司.
Model artifacts used by AtomGradient ASR Runtime for local speech recognition on Intel hardware. The evaluated configuration uses Qwen3-ASR-0.6B with INT8 weight storage and floating-point components. INT8 does not describe every computation in the system.
Model and device specifications
| Item | Specification |
|---|---|
| Base model | Qwen3-ASR-0.6B |
| Model size class | 0.6B |
| Delivered runtime | AtomGradient ASR Runtime, supplied separately from this model repository |
| Evaluated input | 16 kHz, mono, 16-bit PCM |
| Recognition mode | Text returned after processing a complete audio segment |
| Validated device | Intel Core Ultra 5 225H, 32 GiB DDR5 |
| Evaluation focus | Chinese dialogue; public Chinese and English audio samples |
AtomGradient ASR Runtime
AtomGradient ASR Runtime is the local speech-recognition runtime in the AtomGradient speech delivery package. It provides model execution, recognition-request handling, cancellation and dialogue integration, using these model artifacts and third-party inference libraries.
The runtime and the model artifacts are complementary delivery components. This repository distributes the model artifacts; AtomGradient supplies the runtime, installation package and customer deployment documentation separately. The base model and third-party inference libraries retain their respective attribution.
Observed performance
Public audio tests on the device above, 2026-09-11:
| Sample | Audio duration | Processing time |
|---|---|---|
| Chinese | 4.204 s | 0.391 s |
| Short English | 2.461 s | 0.296 s |
| Longer English | 15.051 s | 1.378 s |
Each file was run twice. The table shows the second, warm observation, including audio preprocessing and excluding model loading. These are small-sample measurements, not latency percentiles or a general speed guarantee.
A separate AtomGradient ASR Runtime field test on 2026-09-12 completed 10 of 10 recognition requests, with median local request processing time 0.236 s and range 0.172–1.031 s. This included one wake-phrase verification and nine dialogue requests. Timing excludes the user's speaking time, audio upload and end-of-speech detection. Concurrent speech synthesis can increase recognition latency.
Scope and limitations
- The recorded tests establish that the model runs on the specified device. They do not establish recognition accuracy across arbitrary speakers, languages or environments.
- No held-out WER/CER benchmark or lossless-quantization claim is made. Transcriptions differed between precision variants on the longer English sample.
- The public audio tests and the field-test requests are different inputs with different timing boundaries; they must not be combined into one latency distribution.
- This repository contains model artifacts and associated configuration. AtomGradient ASR Runtime, the installation package and product integration are delivered separately. Model files alone are not the complete demonstrated assistant.
For deployment and integration, contact AtomGradient.
License and attribution
The base model is Qwen/Qwen3-ASR-0.6B. Model artifacts are distributed under Apache-2.0, following the base model's license. AtomGradient's copyright covers its own contributions and does not extend to the upstream Qwen model or third-party components.
- Downloads last month
- 13
Model tree for AtomGradient/Qwen3-ASR-0.6B-int8-ov
Base model
Qwen/Qwen3-ASR-0.6B