Instructions to use micrictor/LFM2.5-350M-ShellAI-v2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use micrictor/LFM2.5-350M-ShellAI-v2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0
Use Docker
docker model run hf.co/micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0
- LM Studio
- Jan
- Ollama
How to use micrictor/LFM2.5-350M-ShellAI-v2-GGUF with Ollama:
ollama run hf.co/micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0
- Unsloth Desktop
- Pi
How to use micrictor/LFM2.5-350M-ShellAI-v2-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use micrictor/LFM2.5-350M-ShellAI-v2-GGUF with Docker Model Runner:
docker model run hf.co/micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0
- Lemonade
How to use micrictor/LFM2.5-350M-ShellAI-v2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0
Run and chat with the model
lemonade run user.LFM2.5-350M-ShellAI-v2-GGUF-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use micrictor/LFM2.5-350M-ShellAI-v2-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use micrictor/LFM2.5-350M-ShellAI-v2-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "micrictor/LFM2.5-350M-ShellAI-v2-GGUF:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
LFM2.5-350M ShellAI v2 โ two-stage Q8_0
This release translates a natural-language shell request through a validated typed intermediate representation before producing Bash:
LFM2.5-350M-ShellAI-v2-IR-Q8_0.gguf: natural language toShellIR-v1.LFM2.5-350M-ShellAI-v2-Command-Q8_0.gguf:ShellIR-v1to one Bash command.
Both files are required for the learned two-stage pipeline. A deterministic ShellIR compiler can replace the second model when minimum memory or latency is more important than retaining the learned lowering stage. The untouched LiquidAI conversational base remains usable without either task adapter.
These are ordinary post-training Q8_0 GGUF files. They are not Liquid AI QAD
checkpoints and are not labeled as QAD.
Evaluation
The sealed suite contains 20 compositional NL-to-Bash cases. No generated command was executed.
| Runtime | Strict / exact | Utility | Valid ShellIR | Gold-IR compiler |
|---|---|---|---|---|
| Transformers BF16 adapter | 80% | 100% | 100% | 100% |
| llama.cpp Q8_0, 1 thread | 75% | 100% | 100% | 100% |
| llama.cpp Q8_0, 2 threads | 75% | 100% | 100% | 100% |
The one Q8_0-only regression was a permission-mask choice in the world-writable-files case. Thread count did not change model output.
| CPU threads | Median IR latency | Median command latency | Median total | IR decode | Command decode |
|---|---|---|---|---|---|
| 1 | 2482 ms | 698 ms | 3525 ms | 35.1 tok/s | 33.0 tok/s |
| 2 | 1381 ms | 368 ms | 1923 ms | 58.1 tok/s | 65.3 tok/s |
Measurements used llama.cpp b10516 with CPU-only inference on the development machine.
See hard_suite_two_stage.json and q8_0_two_stage_gguf.json for per-case results.
Training and selection
The Stage-1 release is an exact LoRA-delta blend of 87.5% safety-focused and 12.5% balanced checkpoints. It improved strict suite accuracy from 70% to 80% while retaining 100% risk/effects-label accuracy. Stage 2 was 100% exact when given gold ShellIR.
Treat model output as untrusted. Validate ShellIR, enforce the risk/effects policy, show destructive commands to the user, and never execute a generated command without explicit authorization.
License
This is a modified derivative of LiquidAI/LFM2.5-350M, distributed under the included
LFM Open License v1.0. See NOTICE for the modification statement.
- Downloads last month
- -
8-bit