Qwen3.8-27B-Splash-Q8 (Compressed 8-bit Baseline)

This repository contains the baseline 8-bit compressed weights for Qwen3.8-27B for the Splash inference engine on Apple Silicon.

Performance

  • Decode Speed: 36.5 tok/s average (peaking at 52.7 tok/s).
  • Speedup: 3.69x over standard autoregressive decoding.
  • Reasoning Accuracy: 44.4% across GPQA Diamond and AIME 2025.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support