about

oq6e quantized version of OxCoder

imatrix used

Default oMLX imatrix

rough performance metrics from quick tests/benchmarks

  • roughly 445 tks/s prefill, 16 tk/s decode on m3 pro 36gb ram in practice on a fresh pi agent
  • context benchmark from omlx for 131k context length started at 2k t/s prefill, finished at ~600 t/s
  • throughput benchmark did not upload because of certain flags used but roughly matched above numbers at 128k
    • ANE prefill with tweaked settings
    • SpecPrefill with qwen3.5 0.8b model

omlx throughput test with ane prefill disabled

https://omlx.ai/benchmarks/performance/tflpzdtq

omlx version used

0.7.0 dev4

Downloads last month
23
Safetensors
Model size
9B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IDFluff/oxcoder-9b-oq6e

Finetuned
Qwen/Qwen3.5-9B
Quantized
(12)
this model