Mellum2 12B A2.5B Thinking - 8-bit MLX

This is an 8-bit affine MLX quantization of JetBrains/Mellum2-12B-A2.5B-Thinking.

Mellum2 Thinking is a reasoning-augmented Mixture-of-Experts assistant model. It has 64 experts with 8 active per token, a 131,072-token context window, and emits reasoning in <think>...</think> blocks before the final answer.

Conversion details

  • Source: JetBrains/Mellum2-12B-A2.5B-Thinking
  • Format: MLX safetensors
  • Quantization: affine, 8 bits, group size 64
  • License: Apache-2.0
  • EOS token: <|im_end|> (token ID 28)

The upstream config.json and generation_config.json identify token ID 0 as the EOS token, while the tokenizer identifies <|im_end|> (ID 28) as EOS. This conversion uses token ID 28 so MLX generation stops at the end of the assistant turn.

Usage

pip install -U mlx-lm

mlx_lm.chat \
  --model mlx-community/Mellum2-12B-A2.5B-Thinking-8bit \
  --max-tokens 8192 \
  --temp 0.6 \
  --top-p 0.95

Benchmarks

Benchmarked with oMLX 0.5.7 and MLX-LM 0.31.3 on an Apple M5 Max MacBook Pro with 128 GB unified memory, using the built-in Python-code workload and 128 generated tokens. Results vary with hardware, runtime versions, context length, and generation settings.

Prompt tokens TTFT Prefill Generation End-to-end Peak memory
4,096 926.4 ms 4,421.3 tok/s 132.3 tok/s 1.9 s 12.7 GB
16,384 3,205.2 ms 5,111.8 tok/s 121.8 tok/s 4.3 s 13.0 GB
32,768 8,171.5 ms 4,010.0 tok/s 112.9 tok/s 9.3 s 13.4 GB
65,536 22,286.9 ms 2,940.6 tok/s 92.9 tok/s 23.7 s 14.1 GB
131,072 67,675.6 ms 1,936.8 tok/s 56.0 tok/s 70.0 s 15.5 GB

At a 4,096-token prompt, batch generation achieved 176.4, 189.3, and 222.8 aggregate generation tok/s at 2, 4, and 8 concurrent requests, respectively, compared with 132.3 tok/s for one request.

Model provenance

For the original model card, training details, benchmark results, and usage guidance, see the upstream JetBrains checkpoint.

Downloads last month
40
Safetensors
Model size
3B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/Mellum2-12B-A2.5B-Thinking-8bit

Quantized
(43)
this model