Qwen3-ASR-0.6B text decoder for the Apple Neural Engine

The text decoder of Qwen/Qwen3-ASR-0.6B, converted to Core ML with ANEMLL 0.3.5.

  • decoder.mlmodelc: 28 layers, two functions over one KV cache (context 512): prefill reads 64 positions per call, infer writes one.
  • lm_head.mlmodelc: final norm is in the decoder; the head returns the best token of each of 16 vocabulary slices (argmax_idx, argmax_val).
  • Weights: LUT8 palettized, 8 channels per group. Needs iOS 18 / macOS 15.

Pair it with the audio encoder and token embeddings from aufklarer/Qwen3-ASR-CoreML (encoder.mlmodelc, embedding.mlmodelc). Audio embeddings go into prefill as hidden states; the chat template and tokenizer are the base model's.

Measured on an M3 Max Neural Engine against the fixed 128-position decoder in aufklarer/Qwen3-ASR-CoreML. Decode times come from live captioning replayed in real time (10 minutes of Mandarin video, each decode stretched 1.1x to match an iPhone 16 Pro). Error rates are against Qwen3-ASR-1.7B's transcript of about 23 minutes of speech.

fixed 128 this build
preview decode, median 506 ms 245 ms
final decode, median 1641 ms 634 ms
decoder busy 62% 32%
character error rate 8.14% 8.19%

License: Apache 2.0, as the base model.

Downloads last month
22
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ukk0708/Qwen3-ASR-0.6B-ANE-decoder

Finetuned
(63)
this model