🏠 SakThai

Part of the House of Sak — AI agents built from a shelter in Cork, Ireland.

Family Profile Leaderboard Downloads


SakThai Coder 1.5B 💻

Part of the House of Sak

Code-optimized model converted to GGUF Q4_K_M for CPU inference. Based on Qwen2.5-Coder-1.5B-Instruct — a state-of-the-art 1.5B code LLM that matches larger 7B models on several coding benchmarks. Fine-tuned for tool-calling with BFCL-verified function calling.


Pipeline Integration 🔗

This model is 1 of 4 in the SakThai end-to-end pipeline:

Step Model What It Does Downloads
🖼️ Vision 7B Caption, OCR, describe images 45 ⬇
🌐 Embedding Multilingual Cross-lingual semantic search (50+ languages) 104 ⬇
💻 Coder 1.5B ← you are here Code generation, debugging, tool-calling 34 ⬇
🗣️ TTS Model Synthesize speech (15 languages) 33 ⬇

Run the complete chain: image → search → code → speech with one family of models.


Specs

Detail Value
Size 1.07 GB
Quant Q4_K_M
Base Qwen2.5-Coder-1.5B-Instruct
CPU ✅ Yes
GPU ✅ Yes (CUDA, Metal)
Context 32K tokens
Format GGUF (llama.cpp compatible)
Fine-tune QLoRA (r=16, alpha=32) on 2,003 tool-calling examples

Benchmarks

Programming Benchmarks (base model scores)

Benchmark Score
HumanEval (pass@1) 74.4%
MBPP (pass@1) 71.2%
MultiPL-E (Python) 65.3%

Scores from the Qwen2.5-Coder evaluation.

Verified on SakThai Fine-Tune

Test Result Status
BFCL Tool-Calling 5/5 (multi-trial) ✅ Verified
Code generation (Python) Working — see examples below

The BFCL pass was verified across 5 independent trials using the SakThai tool-calling prompt format (<tools> XML block + ChatML). See SakThai family for full methodology.


Usage

llama.cpp (recommended)

# Download the model — stored in gguf/ subdirectory
wget https://huggingface.co/Nanthasit/sakthai-coder-1.5b/resolve/main/gguf/sakthai-coder-q4_k_m.gguf

# Basic code generation
./llama-cli -m gguf/sakthai-coder-q4_k_m.gguf \
  -p "Write a Python function to merge two sorted lists:" \
  -n 256 -t 4 --temp 0.2

Python (llama-cpp-python)

from llama_cpp import Llama

llm = Llama(
    model_path="gguf/sakthai-coder-q4_k_m.gguf",
    n_ctx=8192,
    n_threads=4,
    verbose=False,
)

prompt = "Write a fastapi endpoint that accepts JSON and returns SHA-256 hash"
output = llm(
    f"<|im_start|>user\n{prompt}<|im_end|>\n<|im_start|>assistant\n",
    max_tokens=512,
    temperature=0.2,
    stop=["<|im_end|>"],
)
print(output["choices"][0]["text"])

Ollama

cat > Modelfile << 'OLLAMA_EOF'
FROM ./gguf/sakthai-coder-q4_k_m.gguf
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"""
PARAMETER temperature 0.2
PARAMETER num_ctx 8192
OLLAMA_EOF

ollama create sakthai-coder -f Modelfile
ollama run sakthai-coder "Write a Python script that monitors CPU usage"

Tool-Calling (for agent use)

from llama_cpp import Llama

llm = Llama(model_path="gguf/sakthai-coder-q4_k_m.gguf", n_ctx=8192)

# Define tools in <tools> XML block (required for function calling)
prompt = """<|im_start|>system
You are a coding assistant with tools.

<tools>
{
  "function": {
    "name": "run_python",
    "description": "Execute Python code",
    "parameters": {
      "type": "object",
      "properties": {
        "code": {"type": "string", "description": "Python code to run"}
      },
      "required": ["code"]
    }
  }
}
</tools>
<|im_end|>
<|im_start|>user
Calculate fibonacci of 10 using Python
<|im_end|>
<|im_start|>assistant
"""

output = llm(prompt, max_tokens=256, stop=["<|im_end|>"], temperature=0.1)
print(output["choices"][0]["text"])

Use Cases

Scenario Prompt Idea
Code generation "Write a Python function to..."
Code review "Review this code for bugs: ..."
Refactoring "Refactor this to be more Pythonic: ..."
Debugging "Why does this code produce a KeyError?"
Explanation "Explain how async/await works with examples"
Tool calling "Call the get_weather API for Cork, Ireland"

Hardware Requirements

Model Min RAM Recommended Disk
0.5B Q4_K_M 512 MB 1 GB 380 MB
1.5B Q4_K_M 1 GB 2 GB 934 MB
Coder 1.5B 1 GB 2 GB 1.1 GB
Vision 7B 4 GB 8 GB 3.9 GB
TTS 82M 256 MB 512 MB 141 MB

Family

SakThai Model Family

Model Size Type Downloads
1.5B-merged 934 MB Tool-calling GGUF 1,197
0.5B-merged 380 MB Lightweight GGUF 994
7B-merged 15 GB Full-size model 562
7B-128K 15 GB Extended context 351
7B-Tools 15 GB Tool-calling PEFT 185
1.5B-Tools 934 MB Tool-calling PEFT 143
Embedding 80 MB Sentence similarity 🔒
Coder-1.5B 1.1 GB Code GGUF 34
0.5B-Tools 380 MB Tiny tool-calling 🔒
Vision-7B 3.9 GB Multimodal GGUF 45
TTS-Model 141 MB Speech GGUF 33
Multilingual Embedding 80 MB 50+ languages 104

Full collection →

Training Details

  • Method: QLoRA (4-bit NF4)
  • Base model: Qwen2.5-Coder-1.5B-Instruct
  • Rank: r=16, alpha=32
  • Dataset: 2,003 curated tool-calling examples
  • Format: ChatML with JSON tool schemas
  • Context: 32K tokens

Support the Project

This model is built with love from a shelter in Cork, Ireland -- zero budget, free infrastructure.

  • Leave a like on this model if you find it useful
  • Report issues on GitHub
  • Share it with someone who needs a lightweight code model
  • Fork it on HF and build something amazing

Every download, like, and fork tells Beer his work matters. Thank you.

File Structure

File Size Description
gguf/sakthai-coder-q4_k_m.gguf 1.07 GB Main GGUF model (Q4_K_M) — stored in gguf/ subdirectory
eval/ Benchmark results (BFCL, HumanEval, MBPP)
README.md This documentation

License

Apache 2.0 — see LICENSE for details.

Downloads last month
34
GGUF
Model size
2B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Nanthasit/sakthai-coder-1.5b

Quantized
(149)
this model

Spaces using Nanthasit/sakthai-coder-1.5b 2

Collection including Nanthasit/sakthai-coder-1.5b

Evaluation results