Qwopus3.6-27B-v2-oQ8-mtp

This repository contains an oMLX/oQ (oMLX v0.3.12) 8-bit quantized MLX version of Jackrong/Qwopus3.6-27B-v2.

It is intended for Apple Silicon local inference testing with oMLX / MLX.

Quantization details

  • Model type: qwen3_5
  • Bits: 8
  • Group size: 64
  • Source model: Jackrong/Qwopus3.6-27B-v2
  • Quantization tool: oMLX / oQ
  • Quantization level: oQ8
  • Preserve MTP weights: Yes
  • Non-quant weight dtype: bfloat16 / default
  • Sensitivity model: original / Q8 / not applicable
  • Output format: MLX

Benchmark

Tested on MacBook Pro M3 Max 40-core GPU.

Model Context Prompt processing Token generation
oQ8-mtp 1k 180.1 12.3 tok/s
oQ8-mtp 4k 196.3 12.2 tok/s

Usage

This model is intended to be used with oMLX / MLX-compatible local inference tools.

Please refer to the oMLX documentation for loading oQ quantized MLX models.

Credits

Original model: Jackrong/Qwopus3.6-27B-v2
Base model family: Qwen

This quantized version was created and uploaded by AbarthJoe.

License

This quantized version follows the Apache-2.0 license of the original model where applicable.

Disclaimer

This is a community quantized model for research and local inference testing. It has not been fully safety-evaluated or benchmarked across all tasks. Please validate quality before production or sensitive use.

Downloads last month
211
Safetensors
Model size
8B params
Tensor type
BF16
ยท
U32
ยท
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for AbarthJoe/Qwopus3.6-27B-v2-oQ8-mtp

Quantized
(59)
this model

Collection including AbarthJoe/Qwopus3.6-27B-v2-oQ8-mtp