LFM2.5-230M-ToMoE-GGUF

GGUF quantizations of Nichonauta/LFM2.5-230M-ToMoE โ€” the ToMoE Mixture-of-Experts conversion of LiquidAI/LFM2.5-230M.

Files

File Quantization Size BPW
LFM2.5-230M-ToMoE-F16.gguf F16 438 MB 16.0
LFM2.5-230M-ToMoE-Q8_0.gguf Q8_0 233 MB 8.51
LFM2.5-230M-ToMoE-Q4_K_M.gguf Q4_K_M 144 MB 5.26

Important note about the conversion

The ToMoE channel-MoE experts and the per-token dynamic V masks have no native representation in llama.cpp's LFM2 architecture. To produce a runnable GGUF, the pruned model was rebuilt as a dense-equivalent:

  • The cut conv layers were restored to their full width using the original dense weights (the conv weights were never modified by ToMoE โ€” only masked at runtime).
  • The MLP union cut (~99.8% of the FFN) was restored to the full FFN.
  • The attention Q/K were already full-width in the pruned model.

As a result, the GGUF behaves approximately like the dense base model rather than the pruned MoE. For the actual MoE behavior, use the safetensors version with trust_remote_code.

Usage (llama.cpp, CUDA)

llama-server.exe ^
  --model LFM2.5-230M-ToMoE-Q4_K_M.gguf ^
  --gpu-layers all ^
  --ctx-size 32768 ^
  --alias LFM2.5-230M-ToMoE

Or with llama-cli:

llama-cli.exe -m LFM2.5-230M-ToMoE-Q4_K_M.gguf -ngl 99 -p "The capital of France is" -n 25

Measured on an RTX 3060 (CUDA 13.3 build, llama.cpp b10603): ~540โ€“550 t/s generation for Q4_K_M, ~430โ€“450 t/s for Q8_0.

Base model

License

Derivative of LiquidAI/LFM2.5-230M โ€” released under the LFM Open License v1.0 (see LICENSE).

Downloads last month
240
GGUF
Model size
0.2B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Nichonauta/LFM2.5-230M-ToMoE-GGUF

Quantized
(1)
this model