Instructions to use lmoody68/AXIOM-coder-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use lmoody68/AXIOM-coder-9B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf lmoody68/AXIOM-coder-9B:Q4_K_M # Run inference directly in the terminal: llama cli -hf lmoody68/AXIOM-coder-9B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf lmoody68/AXIOM-coder-9B:Q4_K_M # Run inference directly in the terminal: llama cli -hf lmoody68/AXIOM-coder-9B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf lmoody68/AXIOM-coder-9B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf lmoody68/AXIOM-coder-9B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf lmoody68/AXIOM-coder-9B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf lmoody68/AXIOM-coder-9B:Q4_K_M
Use Docker
docker model run hf.co/lmoody68/AXIOM-coder-9B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use lmoody68/AXIOM-coder-9B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "lmoody68/AXIOM-coder-9B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lmoody68/AXIOM-coder-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/lmoody68/AXIOM-coder-9B:Q4_K_M
- Ollama
How to use lmoody68/AXIOM-coder-9B with Ollama:
ollama run hf.co/lmoody68/AXIOM-coder-9B:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use lmoody68/AXIOM-coder-9B with Docker Model Runner:
docker model run hf.co/lmoody68/AXIOM-coder-9B:Q4_K_M
- Lemonade
How to use lmoody68/AXIOM-coder-9B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull lmoody68/AXIOM-coder-9B:Q4_K_M
Run and chat with the model
lemonade run user.AXIOM-coder-9B-Q4_K_M
List all available models
lemonade list
- Atomic Chat
AXIOM-coder-9B
AXIOM is a private, open coding assistant built by Leslie Moody — a Yi-Coder 9B model fine-tuned (QLoRA) on Leslie's own code plus a Python coding-instruction dataset, then merged and quantized to GGUF (Q4_K_M) so it runs fully locally and privately through Ollama or llama.cpp.
It identifies itself as AXIOM, built by Leslie Moody, focuses on coding (writing, explaining, debugging, refactoring), and is tuned to be honest and grounded.
Quick start (Ollama)
- Download
axiom-v1.Q4_K_M.ggufandModelfilefrom this repo into the same folder. - Create and run the model:
ollama create axiom:v1 -f Modelfile
ollama run axiom:v1 "Write a Python function that reverses a linked list, with a docstring."
The included Modelfile carries AXIOM's system prompt (identity + coding focus + honesty) and sensible parameters (temperature 0.25, num_ctx 8192).
llama.cpp
llama-cli -m axiom-v1.Q4_K_M.gguf -p "Refactor this function for readability: ..."
Files
| File | Size | What |
|---|---|---|
axiom-v1.Q4_K_M.gguf |
~5.3 GB | the quantized model (Q4_K_M) |
Modelfile |
tiny | Ollama Modelfile with AXIOM's system prompt + params |
Training
- Base:
01-ai/Yi-Coder-9B-Chat(Apache 2.0) - Method: QLoRA (4-bit), rank 16, 1 epoch
- Data (6,964 examples): ~1,964 docstring↔code pairs extracted from Leslie's own Python projects + 5,000 rows of
iamtarun/python_code_instructions_18k_alpaca - Export: merged to fp16 → converted to GGUF → quantized to Q4_K_M (llama.cpp)
- Trained on a single free-tier GPU.
Intended use
A personal / local coding assistant you fully own and run offline. Great for code help, explanations, and quick refactors.
Limitations (read this)
- It's a 9B fine-tune, not a frontier model. It won't match large hosted assistants on hard, long-context, or agentic coding tasks.
- English + code focused. Non-English ability is limited (the base model isn't broadly multilingual).
- Like all LLMs it can be confident and wrong — verify generated code, and never rely on it for safety- or security-critical decisions without review.
- A small model fine-tuned for 1 epoch can occasionally over-fit phrasing; treat its output as a strong draft, not ground truth.
License & attribution
Released under Apache 2.0, inheriting the base model's license. Built by Leslie Moody. Yi-Coder is by 01.AI. The coding-instruction data is from the dataset linked above.
- Downloads last month
- 8
4-bit