LFM2.5-350M-MLX-8bit

Re-quantization of LiquidAI/LFM2.5-350M (revision 9e6c6ccf47cd318696e137d381a7ded8fe4df09f) with mlx-lm 0.31.3: affine, 8 bit, group size 64, no per-layer overrides — the same recipe as LiquidAI/LFM2.5-350M-MLX-8bit.

The one difference: this checkpoint stores quantization scales and biases in bfloat16, while the official MLX quant stores them in float32. Published so cross-engine benchmarks compare engines under one activation dtype.

Reproduce:

mlx_lm.convert --hf-path LiquidAI/LFM2.5-350M -q --q-bits 8 --q-group-size 64
Downloads last month
-
Safetensors
Model size
0.4B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ipetrukha/LFM2.5-350M-MLX-8bit

Quantized
(75)
this model