Instructions to use AyoubChLin/lfm2.5-2.6b-fable5-coding-agent with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AyoubChLin/lfm2.5-2.6b-fable5-coding-agent with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AyoubChLin/lfm2.5-2.6b-fable5-coding-agent") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("AyoubChLin/lfm2.5-2.6b-fable5-coding-agent") model = AutoModelForCausalLM.from_pretrained("AyoubChLin/lfm2.5-2.6b-fable5-coding-agent", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AyoubChLin/lfm2.5-2.6b-fable5-coding-agent with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AyoubChLin/lfm2.5-2.6b-fable5-coding-agent" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AyoubChLin/lfm2.5-2.6b-fable5-coding-agent", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AyoubChLin/lfm2.5-2.6b-fable5-coding-agent
- SGLang
How to use AyoubChLin/lfm2.5-2.6b-fable5-coding-agent with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AyoubChLin/lfm2.5-2.6b-fable5-coding-agent" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AyoubChLin/lfm2.5-2.6b-fable5-coding-agent", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AyoubChLin/lfm2.5-2.6b-fable5-coding-agent" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AyoubChLin/lfm2.5-2.6b-fable5-coding-agent", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use AyoubChLin/lfm2.5-2.6b-fable5-coding-agent with Docker Model Runner:
docker model run hf.co/AyoubChLin/lfm2.5-2.6b-fable5-coding-agent
LFM2.5-2.6B Fable-5 Coding Agent
This model is a full-parameter supervised fine-tune of LiquidAI/LFM2.5-2.6B on saidutta69/fable-5-premium. It was trained to produce assistant responses, including reasoning-style text and tool-call patterns, from multi-turn chat conversations.
All 2,697,198,592 model parameters were trainable. This is a complete BF16 model checkpoint, not a LoRA, QLoRA, PEFT adapter, or quantized-weight checkpoint. The 8-bit optimizer affected optimizer-state storage only.
Model details
| Field | Value |
|---|---|
| Base model | LiquidAI/LFM2.5-2.6B |
| Architecture | Causal language model |
| Fine-tuning method | Full-parameter supervised fine-tuning |
| Parameters | 2,697,198,592 total; 100% trainable |
| Training precision | BF16, with TF32 enabled |
| Maximum sequence length | 8,192 tokens |
| Training objective | Assistant-only next-token loss |
| Chat formatting | Base model's native chat template |
| Tool-call preprocessing | JSON argument strings converted to mappings for the native LFM2.5 tool-call format |
| Reasoning data | Preserved during training (PRESERVE_THINKING=True) |
The model card does not assign a license. Users should review and comply with the licenses and terms of both the base model and training dataset before use or redistribution.
Intended use
The model is intended for research and evaluation of conversational assistants, coding-agent behavior, long-context supervised fine-tuning, and structured tool-call generation.
Suitable exploratory uses include:
- multi-turn assistant responses;
- code generation and explanation;
- tool-call pattern generation when the application supplies and validates compatible tools; and
- further evaluation or domain adaptation.
The model should not be treated as production-ready on the evidence available here. It has not been evaluated for factuality, security, instruction following, code correctness, tool-call validity, safety, bias, multilingual performance, or robustness.
Training data
The run loaded one Parquet copy of each published openai_chat split to avoid duplicate rows from parallel JSONL and Parquet representations.
| Split | Rows | Mean tokens | P95 tokens | Rows at 8,192-token limit | Mean supervised assistant tokens |
|---|---|---|---|---|---|
| Train | 5,728 | 6,791.6 | 8,192 | 4,520 (78.9%) | 329.5 |
| Validation | 318 | 6,720.3 | 8,192 | 249 (78.3%) | 311.7 |
| Test | 319 | 6,713.9 | 8,192 | 249 (78.1%) | 320.1 |
Only assistant tokens contributed to the loss. System, user, tool-result, and padding tokens were masked with label -100. Assistant tool calls remained supervised. Rows truncated before any assistant output would have been removed; no rows were removed in the recorded run.
Training procedure
| Hyperparameter | Recorded value |
|---|---|
| Epochs | 3 |
| Micro-batch size | 4 |
| Gradient accumulation | 8 |
| Effective batch size | 32 sequences per optimizer step |
| Learning rate | 2e-5 |
| Weight decay | 0.1 |
| Scheduler | Cosine |
| Warmup setting | 0.03 passed through warmup_steps |
| Optimizer | 8-bit AdamW (adamw_bnb_8bit) |
| Gradient clipping | 1.0 |
| Gradient checkpointing | Enabled, non-reentrant |
| Seed / data seed | 42 / 42 |
| Evaluation cadence | Every 100 optimizer steps |
| Checkpoint strategy | Once per epoch, model weights only |
| Hardware | One NVIDIA H200, 139.8 GiB VRAM |
| Software observed | PyTorch 2.8.0+cu129; Transformers 5.15.0 |
The run completed approximately 537 optimizer steps, consistent with 5,728 training examples over three epochs at an effective batch size of 32. The inline comment beside GRADIENT_ACCUMULATION still says the effective batch is 8; the configured values and executed output both establish that the actual effective batch was 32.
Checkpoints were saved with save_only_model=True; they can be evaluated or deployed but cannot exactly resume optimizer and scheduler state.
Results
| Split / metric | Value | Derived perplexity |
|---|---|---|
| Training loss | 0.7474 | 2.1115 |
| Validation loss | 0.3403 | 1.4053 |
| Held-out test loss | 0.3388 | 1.4033 |
Training completed in 9,400.4 seconds (about 2 hours 36 minutes 40 seconds), at 1.828 samples/second and 0.057 optimizer steps/second. The run reported approximately 2.05e18 floating-point operations.
Perplexity is calculated as exp(loss). These losses cover only the assistant tokens selected by the masking procedure, so they are not directly comparable with full-sequence language-model losses. The training loss is averaged over the optimization trajectory, while validation and test losses were measured after training; their values should not be compared as if they were measured at the same checkpoint under identical conditions.
No pre-fine-tuning baseline, external benchmark, confidence interval, or repeated-seed result was recorded. Consequently, the results establish held-out token-loss performance for this run but do not by themselves demonstrate an improvement over the base model or general task quality.
Qualitative observation
On a prompt requesting a Python interval-merging function, the model produced a reasonable high-level plan and emitted a Write(...) tool call despite no explicit tool schema being supplied in the prompt. The generated program was not correct as written: one multiline assert was syntactically invalid, and an adjacency test contradicted the implementation's last_end + 1 merge rule.
This single example is illustrative, not an evaluation. It highlights three deployment requirements:
- Do not expose reasoning traces when the product requires hidden internal reasoning.
- Parse, authorize, and validate every tool call in a sandbox; never execute generated calls directly.
- Execute generated code against tests; plausible structure and self-generated assertions do not establish correctness.
Inference
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "AyoubChLin/lfm2.5-2.6b-fable5-coding-agent"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
dtype=torch.bfloat16,
device_map="auto",
)
messages = [
{"role": "system", "content": "You are a careful coding assistant."},
{"role": "user", "content": "Write a tested Python function that merges overlapping intervals."},
]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
).to(model.device)
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=768,
do_sample=True,
temperature=0.2,
top_k=50,
repetition_penalty=1.1,
pad_token_id=tokenizer.pad_token_id,
eos_token_id=tokenizer.eos_token_id,
)
new_tokens = output[0, inputs["input_ids"].shape[1]:]
print(tokenizer.decode(new_tokens, skip_special_tokens=False))
Use the base model's native chat template. If tool use is enabled, provide explicit tool definitions in the serving layer and validate generated calls before execution.
Reproducibility notes
The source run used an NVIDIA H200 with BF16 support. It normalized OpenAI-style tool calls, tokenized with the native LFM2.5 chat template, preserved thinking content, and trained on assistant tokens only. The held-out test split was not used for optimization or periodic validation.
The training notebook is the authoritative source for implementation details. Before reproducing the run, import TrainingArguments before inspecting its signature, correct the stale effective-batch comment, restart the kernel, and execute every cell in order so the source and outputs cannot diverge.
- Downloads last month
- 397