Qwen3.8-27B-Splash-Q8 (Compressed 8-bit Baseline)
This repository contains the baseline 8-bit compressed weights for Qwen3.8-27B for the Splash inference engine on Apple Silicon.
Performance
- Decode Speed: 36.5 tok/s average (peaking at 52.7 tok/s).
- Speedup: 3.69x over standard autoregressive decoding.
- Reasoning Accuracy: 44.4% across GPQA Diamond and AIME 2025.