Instructions to use AbarthJoe/Qwopus3.6-27B-v2-oQ8-mtp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use AbarthJoe/Qwopus3.6-27B-v2-oQ8-mtp with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Qwopus3.6-27B-v2-oQ8-mtp AbarthJoe/Qwopus3.6-27B-v2-oQ8-mtp
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
Qwopus3.6-27B-v2-oQ8-mtp
This repository contains an oMLX/oQ (oMLX v0.3.12) 8-bit quantized MLX version of Jackrong/Qwopus3.6-27B-v2.
It is intended for Apple Silicon local inference testing with oMLX / MLX.
Quantization details
- Model type: qwen3_5
- Bits: 8
- Group size: 64
- Source model: Jackrong/Qwopus3.6-27B-v2
- Quantization tool: oMLX / oQ
- Quantization level: oQ8
- Preserve MTP weights: Yes
- Non-quant weight dtype: bfloat16 / default
- Sensitivity model: original / Q8 / not applicable
- Output format: MLX
Benchmark
Tested on MacBook Pro M3 Max 40-core GPU.
| Model | Context | Prompt processing | Token generation |
|---|---|---|---|
| oQ8-mtp | 1k | 180.1 | 12.3 tok/s |
| oQ8-mtp | 4k | 196.3 | 12.2 tok/s |
Usage
This model is intended to be used with oMLX / MLX-compatible local inference tools.
Please refer to the oMLX documentation for loading oQ quantized MLX models.
Credits
Original model: Jackrong/Qwopus3.6-27B-v2
Base model family: Qwen
This quantized version was created and uploaded by AbarthJoe.
License
This quantized version follows the Apache-2.0 license of the original model where applicable.
Disclaimer
This is a community quantized model for research and local inference testing. It has not been fully safety-evaluated or benchmarked across all tasks. Please validate quality before production or sensitive use.
- Downloads last month
- 211
8-bit
Model tree for AbarthJoe/Qwopus3.6-27B-v2-oQ8-mtp
Base model
Jackrong/Qwopus3.6-27B-v2