Laguna-XS-2.1-oQ2

This model was quantized using oQ (oMLX v0.5.4) mixed-precision quantization.

Base model: poolside/Laguna-XS-2.1

Quantization details

  • Model type: laguna
  • Bits: 2
  • Group size: 64
  • Format: MLX safetensors

Environment

  • Hardware: M5 MacBook Air 32GB
  • Inference Framework: oMLX v0.5.4
  • Max Concurrent Requests: 4
  • Settings:
    • Thinking: Disabled
    • TurboQuant KV Cache: Enabled

Performance Benchmarks

Note: Results are for reference only and may vary depending on hardware, software configuration, and workload.

Single Request Results

Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem
pp1024/tg128 1201.1 20.47 853.4 tok/s 49.2 tok/s 3.812 302.5 tok/s 11.51 GB
pp4096/tg128 5103.7 22.63 802.8 tok/s 44.5 tok/s 7.987 529.0 tok/s 11.64 GB

Continuous Batching (pp1024 / tg128)

Batch tg TPS Speedup pp TPS pp TPS/req TTFT(ms) E2E(s)
1x 49.2 tok/s 1.00x 853.4 tok/s 853.4 tok/s 1201.1 3.812
2x 65.1 tok/s 1.32x 764.1 tok/s 382.1 tok/s 2680.3 6.615
4x 94.6 tok/s 1.92x 1226.5 tok/s 306.6 tok/s 3254.9 8.750

Intelligence Benchmark

Note: Each benchmark round tests only 30 questions. Results are for reference only.

Benchmark Accuracy Correct Total Time(s) Think
MMLU 30.0% 9 30 74.6 No
TRUTHFULQA 56.7% 17 30 55.6 No
GSM8K 86.7% 26 30 134.7 No
MATHQA 20.0% 6 30 97 No
HUMANEVAL 76.7% 23 30 106.7 No
Downloads last month
104
Safetensors
Model size
3B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support