Qwen3.8 27B Core AI 128K

Core AI export of Qwen3.8-27B for macOS with INT4 linear weights and a 131,072-token dynamic KV-cache bound. The bundle is a single-token GPU-pipelined decode graph; prompt prefill is pipelined token-wise by the host runtime. It is intended for Apple Silicon running the Core AI runtime in macOS 27 or newer.

This is an independent conversion of Qwen/Qwen3.8-27B at revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0.

Layout and verification

The language bundle is at gpu-pipelined/qwen3_8_27b_decode_int4lin. SHA256SUMS covers every published file. The compiled payload is 18,791,357,252 bytes with SHA-256 6aeac000c7073f4382a0fa2dc871ec1df65aaf82d04f7d9b0d36fb1e2b420a51.

On an Apple M4 Pro with 48 GB unified memory, a 128-token prompt and 256-token decode measured approximately 10.1 prompt tokens/s and 10.0 generated tokens/s. Performance depends on hardware, OS, prompt length, memory pressure, and thermal state.

License

The converted model remains subject to the Apache License 2.0 supplied with the source model. See LICENSE. This conversion is independent and is not produced or endorsed by Qwen.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ETeissonniere/Qwen3.8-27B-CoreAI-128K

Base model

Qwen/Qwen3.8-27B
Finetuned
(259)
this model