Instructions to use Nanthasit/sakthai-coder-1.5b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Nanthasit/sakthai-coder-1.5b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Nanthasit/sakthai-coder-1.5b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Nanthasit/sakthai-coder-1.5b", device_map="auto") - llama-cpp-python
How to use Nanthasit/sakthai-coder-1.5b with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="Nanthasit/sakthai-coder-1.5b", filename="gguf/sakthai-coder-q4_k_m.gguf", )
llm.create_chat_completion( messages = [ { "role": "user", "content": "What is the capital of France?" } ] ) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Nanthasit/sakthai-coder-1.5b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M # Run inference directly in the terminal: llama cli -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M # Run inference directly in the terminal: llama cli -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M
Use Docker
docker model run hf.co/Nanthasit/sakthai-coder-1.5b:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Nanthasit/sakthai-coder-1.5b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Nanthasit/sakthai-coder-1.5b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nanthasit/sakthai-coder-1.5b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Nanthasit/sakthai-coder-1.5b:Q4_K_M
- SGLang
How to use Nanthasit/sakthai-coder-1.5b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Nanthasit/sakthai-coder-1.5b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nanthasit/sakthai-coder-1.5b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Nanthasit/sakthai-coder-1.5b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nanthasit/sakthai-coder-1.5b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use Nanthasit/sakthai-coder-1.5b with Ollama:
ollama run hf.co/Nanthasit/sakthai-coder-1.5b:Q4_K_M
- Unsloth Studio
How to use Nanthasit/sakthai-coder-1.5b with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Nanthasit/sakthai-coder-1.5b to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Nanthasit/sakthai-coder-1.5b to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Nanthasit/sakthai-coder-1.5b to start chatting
- Pi
How to use Nanthasit/sakthai-coder-1.5b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Nanthasit/sakthai-coder-1.5b:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use Nanthasit/sakthai-coder-1.5b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Nanthasit/sakthai-coder-1.5b:Q4_K_M
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use Nanthasit/sakthai-coder-1.5b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Nanthasit/sakthai-coder-1.5b:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use Nanthasit/sakthai-coder-1.5b with Docker Model Runner:
docker model run hf.co/Nanthasit/sakthai-coder-1.5b:Q4_K_M
- Lemonade
How to use Nanthasit/sakthai-coder-1.5b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Nanthasit/sakthai-coder-1.5b:Q4_K_M
Run and chat with the model
lemonade run user.sakthai-coder-1.5b-Q4_K_M
List all available models
lemonade list
SakThai Coder 1.5B 💻
Part of the House of Sak
Code-optimized model converted to GGUF Q4_K_M for CPU inference. Based on Qwen2.5-Coder-1.5B-Instruct — a state-of-the-art 1.5B code LLM that matches larger 7B models on several coding benchmarks. Fine-tuned for tool-calling with BFCL-verified function calling.
Pipeline Integration 🔗
This model is 1 of 4 in the SakThai end-to-end pipeline:
| Step | Model | What It Does | Downloads |
|---|---|---|---|
| 🖼️ | Vision 7B | Caption, OCR, describe images | 45 ⬇ |
| 🌐 | Embedding Multilingual | Cross-lingual semantic search (50+ languages) | 104 ⬇ |
| 💻 | Coder 1.5B ← you are here | Code generation, debugging, tool-calling | 34 ⬇ |
| 🗣️ | TTS Model | Synthesize speech (15 languages) | 33 ⬇ |
Run the complete chain: image → search → code → speech with one family of models.
Specs
| Detail | Value |
|---|---|
| Size | 1.07 GB |
| Quant | Q4_K_M |
| Base | Qwen2.5-Coder-1.5B-Instruct |
| CPU | ✅ Yes |
| GPU | ✅ Yes (CUDA, Metal) |
| Context | 32K tokens |
| Format | GGUF (llama.cpp compatible) |
| Fine-tune | QLoRA (r=16, alpha=32) on 2,003 tool-calling examples |
Benchmarks
Programming Benchmarks (base model scores)
| Benchmark | Score |
|---|---|
| HumanEval (pass@1) | 74.4% |
| MBPP (pass@1) | 71.2% |
| MultiPL-E (Python) | 65.3% |
Scores from the Qwen2.5-Coder evaluation.
Verified on SakThai Fine-Tune
| Test | Result | Status |
|---|---|---|
| BFCL Tool-Calling | 5/5 (multi-trial) | ✅ Verified |
| Code generation (Python) | Working — see examples below | ✅ |
The BFCL pass was verified across 5 independent trials using the SakThai tool-calling prompt format (<tools> XML block + ChatML). See SakThai family for full methodology.
Usage
llama.cpp (recommended)
# Download the model — stored in gguf/ subdirectory
wget https://huggingface.co/Nanthasit/sakthai-coder-1.5b/resolve/main/gguf/sakthai-coder-q4_k_m.gguf
# Basic code generation
./llama-cli -m gguf/sakthai-coder-q4_k_m.gguf \
-p "Write a Python function to merge two sorted lists:" \
-n 256 -t 4 --temp 0.2
Python (llama-cpp-python)
from llama_cpp import Llama
llm = Llama(
model_path="gguf/sakthai-coder-q4_k_m.gguf",
n_ctx=8192,
n_threads=4,
verbose=False,
)
prompt = "Write a fastapi endpoint that accepts JSON and returns SHA-256 hash"
output = llm(
f"<|im_start|>user\n{prompt}<|im_end|>\n<|im_start|>assistant\n",
max_tokens=512,
temperature=0.2,
stop=["<|im_end|>"],
)
print(output["choices"][0]["text"])
Ollama
cat > Modelfile << 'OLLAMA_EOF'
FROM ./gguf/sakthai-coder-q4_k_m.gguf
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"""
PARAMETER temperature 0.2
PARAMETER num_ctx 8192
OLLAMA_EOF
ollama create sakthai-coder -f Modelfile
ollama run sakthai-coder "Write a Python script that monitors CPU usage"
Tool-Calling (for agent use)
from llama_cpp import Llama
llm = Llama(model_path="gguf/sakthai-coder-q4_k_m.gguf", n_ctx=8192)
# Define tools in <tools> XML block (required for function calling)
prompt = """<|im_start|>system
You are a coding assistant with tools.
<tools>
{
"function": {
"name": "run_python",
"description": "Execute Python code",
"parameters": {
"type": "object",
"properties": {
"code": {"type": "string", "description": "Python code to run"}
},
"required": ["code"]
}
}
}
</tools>
<|im_end|>
<|im_start|>user
Calculate fibonacci of 10 using Python
<|im_end|>
<|im_start|>assistant
"""
output = llm(prompt, max_tokens=256, stop=["<|im_end|>"], temperature=0.1)
print(output["choices"][0]["text"])
Use Cases
| Scenario | Prompt Idea |
|---|---|
| Code generation | "Write a Python function to..." |
| Code review | "Review this code for bugs: ..." |
| Refactoring | "Refactor this to be more Pythonic: ..." |
| Debugging | "Why does this code produce a KeyError?" |
| Explanation | "Explain how async/await works with examples" |
| Tool calling | "Call the get_weather API for Cork, Ireland" |
Hardware Requirements
| Model | Min RAM | Recommended | Disk |
|---|---|---|---|
| 0.5B Q4_K_M | 512 MB | 1 GB | 380 MB |
| 1.5B Q4_K_M | 1 GB | 2 GB | 934 MB |
| Coder 1.5B | 1 GB | 2 GB | 1.1 GB |
| Vision 7B | 4 GB | 8 GB | 3.9 GB |
| TTS 82M | 256 MB | 512 MB | 141 MB |
Family
- 1.5B SakThai (general) — tool-calling & general chat
- 0.5B SakThai (lightweight) — ultra-lightweight GGUF
- 7B SakThai (full-size) — high-capacity model
- 1.5B Coder (this one) — code generation GGUF
SakThai Model Family
| Model | Size | Type | Downloads |
|---|---|---|---|
| 1.5B-merged | 934 MB | Tool-calling GGUF | 1,197 |
| 0.5B-merged | 380 MB | Lightweight GGUF | 994 |
| 7B-merged | 15 GB | Full-size model | 562 |
| 7B-128K | 15 GB | Extended context | 351 |
| 7B-Tools | 15 GB | Tool-calling PEFT | 185 |
| 1.5B-Tools | 934 MB | Tool-calling PEFT | 143 |
| Embedding | 80 MB | Sentence similarity | 🔒 |
| Coder-1.5B | 1.1 GB | Code GGUF | 34 |
| 0.5B-Tools | 380 MB | Tiny tool-calling | 🔒 |
| Vision-7B | 3.9 GB | Multimodal GGUF | 45 |
| TTS-Model | 141 MB | Speech GGUF | 33 |
| Multilingual Embedding | 80 MB | 50+ languages | 104 |
Training Details
- Method: QLoRA (4-bit NF4)
- Base model: Qwen2.5-Coder-1.5B-Instruct
- Rank: r=16, alpha=32
- Dataset: 2,003 curated tool-calling examples
- Format: ChatML with JSON tool schemas
- Context: 32K tokens
Support the Project
This model is built with love from a shelter in Cork, Ireland -- zero budget, free infrastructure.
- Leave a like on this model if you find it useful
- Report issues on GitHub
- Share it with someone who needs a lightweight code model
- Fork it on HF and build something amazing
Every download, like, and fork tells Beer his work matters. Thank you.
File Structure
| File | Size | Description |
|---|---|---|
| gguf/sakthai-coder-q4_k_m.gguf | 1.07 GB | Main GGUF model (Q4_K_M) — stored in gguf/ subdirectory |
| eval/ | — | Benchmark results (BFCL, HumanEval, MBPP) |
| README.md | — | This documentation |
License
Apache 2.0 — see LICENSE for details.
- Downloads last month
- 34
4-bit
Model tree for Nanthasit/sakthai-coder-1.5b
Base model
Qwen/Qwen2.5-1.5BSpaces using Nanthasit/sakthai-coder-1.5b 2
Collection including Nanthasit/sakthai-coder-1.5b
Evaluation results
- Pass@5 on BFCLself-reported5/5
- pass@1 on HumanEvalself-reported74.400
- pass@1 on MBPPself-reported71.200