LFM2.5-8B-A1B-UltraCoder-L3

This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.

These artifacts are compatible with Hugging Face Transformers and llama.cpp (GGUF).

A Q4_K_M GGUF variant is provided for resource-efficient local inference using llama.cpp-compatible runtimes.

LFM2.5-8B-A1B-UltraCoder-L3 is an experimental coding-specialized derivative of LiquidAI's LFM2.5-8B-A1B family.

Built on the architectural foundation of LiquidAI/LFM2.5-8B-A1B-Base, the model preserves the original LFM2 sparse Mixture-of-Experts (MoE) architecture while specializing its behavior toward Python programming, algorithmic problem-solving, code generation, and coding-assistant use cases.

UltraCoder Highlights

The complete adaptation used approximately 55.30M input tokens in two stages:

  1. L2 continued pretraining (CPT) on high-quality Python source code from openbmb/UltraData-Code.
  2. L3 supervised fine-tuning (SFT) on coding tasks containing problem statements, reasoning/analysis, and solutions.

Model Overview

  • Type: Sparse Mixture-of-Experts Causal Language Model
  • Training Stage: Continued Pre-training & Supervised Fine-Tuning
  • Number of Parameters: ~8.3B
  • Active Parameters per Token: ~1.5B
  • Number of Layers: 24
  • Number of Experts: 32 (4 activated per token)
  • Context Length: 8,192 tokens (upstream capability 131,072)

Benchmark Results

Under a controlled Q4_K_M EvalPlus comparison against the Base LFM2.5-8B-A1B Q4_K_M model, UltraCoder demonstrated substantial improvements across coding benchmarks.

Text Performance

UltraCoder Q4_K_M Base LFM2.5 Q4_K_M Delta pp Relative Gain
EvalPlus Benchmarks
HumanEval
58.54% 41.46% +17.07 pp +41.18%
HumanEval+
53.05% 39.63% +13.41 pp +33.85%
MBPP
61.38% 53.44% +7.94 pp +14.85%
MBPP+
49.21% 46.83% +2.38 pp +5.08%
Coding+ Avg
51.13% 43.23% +7.90 pp +18.27%

Quickstart

llama.cpp — Q4_K_M

With a recent llama.cpp installation:

llama serve \
  -hf Susant-Achary/LFM2.5-8B-A1B-UltraCoder-L3:Q4_K_M \
  -c 8192

For CLI inference:

llama cli \
  -hf Susant-Achary/LFM2.5-8B-A1B-UltraCoder-L3:Q4_K_M \
  -c 8192

Transformers

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "Susant-Achary/LFM2.5-8B-A1B-UltraCoder-L3"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    dtype=torch.bfloat16,
    device_map="auto",
)

messages = [
    {
        "role": "user",
        "content": "Write an efficient Python implementation of Dijkstra's shortest-path algorithm."
    }
]

inputs = tokenizer.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_tensors="pt",
).to(model.device)

outputs = model.generate(inputs, max_new_tokens=1024, do_sample=False)
generated = outputs[0, inputs.shape[-1]:]
print(tokenizer.decode(generated, skip_special_tokens=True))
Downloads last month
-
Safetensors
Model size
8B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Susant-Achary/LFM2.5-8B-A1B-UltraCoder-L3

Quantized
(92)
this model
Quantizations
2 models

Dataset used to train Susant-Achary/LFM2.5-8B-A1B-UltraCoder-L3