Qwen3-ASR-0.6B for Intel Core Ultra

Maintained by AtomGradient / 质子梯度(北京)科技有限公司.

Model artifacts used by AtomGradient ASR Runtime for local speech recognition on Intel hardware. The evaluated configuration uses Qwen3-ASR-0.6B with INT8 weight storage and floating-point components. INT8 does not describe every computation in the system.

Model and device specifications

Item Specification
Base model Qwen3-ASR-0.6B
Model size class 0.6B
Delivered runtime AtomGradient ASR Runtime, supplied separately from this model repository
Evaluated input 16 kHz, mono, 16-bit PCM
Recognition mode Text returned after processing a complete audio segment
Validated device Intel Core Ultra 5 225H, 32 GiB DDR5
Evaluation focus Chinese dialogue; public Chinese and English audio samples

AtomGradient ASR Runtime

AtomGradient ASR Runtime is the local speech-recognition runtime in the AtomGradient speech delivery package. It provides model execution, recognition-request handling, cancellation and dialogue integration, using these model artifacts and third-party inference libraries.

The runtime and the model artifacts are complementary delivery components. This repository distributes the model artifacts; AtomGradient supplies the runtime, installation package and customer deployment documentation separately. The base model and third-party inference libraries retain their respective attribution.

Observed performance

Public audio tests on the device above, 2026-09-11:

Sample Audio duration Processing time
Chinese 4.204 s 0.391 s
Short English 2.461 s 0.296 s
Longer English 15.051 s 1.378 s

Each file was run twice. The table shows the second, warm observation, including audio preprocessing and excluding model loading. These are small-sample measurements, not latency percentiles or a general speed guarantee.

A separate AtomGradient ASR Runtime field test on 2026-09-12 completed 10 of 10 recognition requests, with median local request processing time 0.236 s and range 0.172–1.031 s. This included one wake-phrase verification and nine dialogue requests. Timing excludes the user's speaking time, audio upload and end-of-speech detection. Concurrent speech synthesis can increase recognition latency.

Scope and limitations

  • The recorded tests establish that the model runs on the specified device. They do not establish recognition accuracy across arbitrary speakers, languages or environments.
  • No held-out WER/CER benchmark or lossless-quantization claim is made. Transcriptions differed between precision variants on the longer English sample.
  • The public audio tests and the field-test requests are different inputs with different timing boundaries; they must not be combined into one latency distribution.
  • This repository contains model artifacts and associated configuration. AtomGradient ASR Runtime, the installation package and product integration are delivered separately. Model files alone are not the complete demonstrated assistant.

For deployment and integration, contact AtomGradient.

License and attribution

The base model is Qwen/Qwen3-ASR-0.6B. Model artifacts are distributed under Apache-2.0, following the base model's license. AtomGradient's copyright covers its own contributions and does not extend to the upstream Qwen model or third-party components.

Downloads last month
13
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AtomGradient/Qwen3-ASR-0.6B-int8-ov

Quantized
(52)
this model