Mellum2 12B A2.5B Instruct - 4-bit MLX

This is a 4-bit affine MLX quantization of JetBrains/Mellum2-12B-A2.5B-Instruct.

Mellum2 Instruct is a Mixture-of-Experts assistant model with 64 experts and 8 active experts per token. It supports a 131,072-token context window and is optimized for direct instruction following.

Conversion details

  • Source: JetBrains/Mellum2-12B-A2.5B-Instruct
  • Format: MLX safetensors
  • Quantization: affine, 4 bits, group size 64
  • License: Apache-2.0
  • EOS token: <|im_end|> (token ID 28)

The upstream config.json and generation_config.json identify token ID 0 as the EOS token, while the tokenizer identifies <|im_end|> (ID 28) as EOS. This conversion uses token ID 28 so MLX generation stops at the end of the assistant turn.

Usage

pip install -U mlx-lm

mlx_lm.chat \
  --model mlx-community/Mellum2-12B-A2.5B-Instruct-4bit \
  --max-tokens 8192 \
  --temp 0.6 \
  --top-p 0.95

Model provenance

For the original model card, training details, benchmark results, and usage guidance, see the upstream JetBrains checkpoint.

Downloads last month
14
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/Mellum2-12B-A2.5B-Instruct-4bit

Quantized
(23)
this model