Instructions to use kimjg/toolRouter-1B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use kimjg/toolRouter-1B with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/llama-3.2-1b-instruct-unsloth-bnb-4bit") model = PeftModel.from_pretrained(base_model, "kimjg/toolRouter-1B") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use kimjg/toolRouter-1B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf kimjg/toolRouter-1B:Q4_K_M # Run inference directly in the terminal: llama cli -hf kimjg/toolRouter-1B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf kimjg/toolRouter-1B:Q4_K_M # Run inference directly in the terminal: llama cli -hf kimjg/toolRouter-1B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf kimjg/toolRouter-1B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf kimjg/toolRouter-1B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf kimjg/toolRouter-1B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf kimjg/toolRouter-1B:Q4_K_M
Use Docker
docker model run hf.co/kimjg/toolRouter-1B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use kimjg/toolRouter-1B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "kimjg/toolRouter-1B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kimjg/toolRouter-1B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/kimjg/toolRouter-1B:Q4_K_M
- Ollama
How to use kimjg/toolRouter-1B with Ollama:
ollama run hf.co/kimjg/toolRouter-1B:Q4_K_M
- Unsloth Desktop
- Pi
How to use kimjg/toolRouter-1B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kimjg/toolRouter-1B:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "kimjg/toolRouter-1B:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use kimjg/toolRouter-1B with Docker Model Runner:
docker model run hf.co/kimjg/toolRouter-1B:Q4_K_M
- Lemonade
How to use kimjg/toolRouter-1B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull kimjg/toolRouter-1B:Q4_K_M
Run and chat with the model
lemonade run user.toolRouter-1B-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use kimjg/toolRouter-1B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kimjg/toolRouter-1B:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default kimjg/toolRouter-1B:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use kimjg/toolRouter-1B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kimjg/toolRouter-1B:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "kimjg/toolRouter-1B:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
toolRouter-1B
ํ๊ตญ์ด ์ฐ์ ํด์ฝ ๋ผ์ฐํฐ โ ์ฌ์ฉ์ ๋ฐํ๋ฅผ ๋ฐ์ ํญ์ ํด์ฝ ํ๋๋ก ์๋ตํ๋ ๋ผ์ฐํ ๋ชจ๋ธ์ ๋๋ค. Llama-3.2-1B-Instruct ์์ ์น๋ QLoRA ์ด๋ํฐ(~45MB)๋ก, 4bit ๋ก๋ ์ VRAM ์ฝ 2.5GB, GGUF Q4 ๋ณํ ์ ์ฝ 1GB๋ก ๊ตฌ๋๋ฉ๋๋ค.
- ์์ ํ
์คํธ๋ฅผ ์์ฑํ์ง ์์ต๋๋ค. ๋ชจ๋ ์ถ๋ ฅ์
<tool_call>{"name": ..., "arguments": {...}}</tool_call>ํ์์ ๋๋ค. - ์ผ๋ฐ ๋ํยท์์ ์๋ต์
reply(text=...), ์ฒ๋ฆฌ ๋ถ๊ฐ ์์ฒญ์escalate(query=์๋ฌธ)ํด๋ก ํํํฉ๋๋ค. - ์ธ์๋ ์ฌ์ฉ์ ๋ฌธ์ฅ์ ํํ์ ์๋ฌธ ๊ทธ๋๋ก ์ถ์ถํ๋๋ก ํ์ต๋์ต๋๋ค (ํ๊ตญ์ด ์ธ์์ ์์ด ๋ฒ์ญยท์์ฝยท์ง์ด๋ ์ต์ ).
- ์ฒ์ ๋ณด๋ ํด ์คํค๋ง(zero-shot toolset)์์๋ ๋์ํ๋๋ก ํ์ต๋์ต๋๋ค.
- ๋ฉํฐํด์ด ํ์ํ๋ฉด recallResolver-1B์ ์ง์ผ๋ก ์ฐ์ธ์ โ ๊ฐ์ ๋ฒ ์ด์ค๋ฅผ ๊ณต์ ํด ์ด๋ํฐ ์ ํ๋ง์ผ๋ก ํจ๊ป ๋์ํ๋ฉฐ, ๋ชจ๋ธ์ ๋ฐ๋ก ๋์ฐ๋ ๊ฒ ๋๋น VRAM์ด ์ ๋ฐ์ ๋๋ค โ ์ค์ธก: llama.cpp ๋ฒ ์ด์ค Q4 + LoRA 2๊ฐ ํ ์๋ฒ = 1.5GB, ๋จธ์ง Q4 ์๋ฒ 2๊ฐ = 2.8GB (๊ฐ 2048 ctx, ์ถ๋ก ๋ฒํผ ํฌํจ).
์ต์ ๋ฒ์ โ 2026-09.2 (recall ์ง์)
ํ๊ตญ์ด ์ค์ฌ์ฉ ๋ฌธ์ฒด 6,000๋ฌธ์ฅ ์ง์ ์์ฑ ๋ฐ์ดํฐ๋ก ๊ฐํํ๊ณ , **๊ณผ๊ฑฐ ๋ํ ์ฐธ์กฐ ๊ฐ์ง(recall)**๋ฅผ ์ถ๊ฐํ ๋ฒ์ ์
๋๋ค.
| ํ๊ฐ | ์์น |
|---|---|
| ํ๊ตญ์ด ํ๊ฐ์ 396 (decision / tool / args) | 97.2 / 95.7 / 83.8 |
| held-out 1,100 (decision / tool / args) | 99.5 / 99.2 / 97.0 |
| BFCL v3 single-turn (simple / multiple / irrelevance / overall) | 86.0 / 85.5 / 80.0 / 84.2 |
| recall ํ์ (๊ฒฝ๊ณ ์ผ์ด์ค ํฌํจ 120) | ์ฌํ 86.7% / ์ ๋ฐ 91.2% / ํจ์ ๋ฐฉ์ด 52/55 |
์ฌ์ฉ๋ฒ
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = "meta-llama/Llama-3.2-1B-Instruct"
model = AutoModelForCausalLM.from_pretrained(base, device_map="auto")
model = PeftModel.from_pretrained(model, "kimjg/toolRouter-1B")
tokenizer = AutoTokenizer.from_pretrained("kimjg/toolRouter-1B")
์ฑํ
ํ
ํ๋ฆฟ์ ์ ์ฅ์์ chat_template.jinja์ ํฌํจ๋ผ ์์ต๋๋ค. greedy ๋์ฝ๋ฉ์ ๊ถ์ฅํ๋ฉฐ, ํ๋ก๋์
์์๋ xgrammar ๋ฑ JSON ์คํค๋ง ๋ฌธ๋ฒ ๊ฐ์ (constrained decoding)๋ฅผ ํจ๊ป ์ฐ๋ฉด ํ์ฑยท์คํค๋ง ํต๊ณผ์จ 100%๋ฅผ ๋ณด์ฅํ ์ ์์ต๋๋ค.
ํด ์ง์ ๋ฐฉ๋ฒ
์์คํ ํ๋กฌํํธ์ ํด ์คํค๋ง๋ฅผ ํ ์ค์ ํ๋์ฉ compact JSON์ผ๋ก ๋์ดํฉ๋๋ค. ํ์ต ๋ ์ฌ์ฉํ ํ ํ๋ฆฟ ๊ทธ๋๋ก ์ฐ๋ ๊ฒ์ ๊ถ์ฅํฉ๋๋ค:
๋๋ ๊ฒ์ ๊ณต๋ต ๋์ฐ๋ฏธ์ ํด์ฝ ๋ผ์ฐํฐ๋ค. ์ฌ์ฉ์์ ์์ฒญ์ ์ฝ๊ณ ์๋ ํด ์ค ์ ํํ ํ๋๋ฅผ ํธ์ถํ๋ ๊ฒ์ผ๋ก๋ง ์๋ตํ๋ค. ์์ ํ
์คํธ๋ฅผ ์ถ๋ ฅํ์ง ์๋๋ค.
- ์ผ๋ฐ ๋ํ๋ ์ง์ ๋ตํ ์ ์๋ ์ง๋ฌธ์ reply ํด์ ํธ์ถํ๋ค.
- ๋ชฉ๋ก์ ์๋ ํด๋ก ์ฒ๋ฆฌํ ์ ์๊ฑฐ๋ ๋ฒ์๋ฅผ ๋ฒ์ด๋ ์์ฒญ์ escalate ํด์ ํธ์ถํ๋ค.
- ํด ์ธ์๋ ์ฌ์ฉ์์ ์์ฒญ์์ ์ถ์ถํ๋ค. ์์ฒญ์ ์๋ ๊ฐ์ ์ง์ด๋ด์ง ์๋๋ค.
์๋ต ํ์์ ๋ฐ๋์ ๋ค์๊ณผ ๊ฐ๋ค:
<tool_call>{"name": "<ํด ์ด๋ฆ>", "arguments": {<์ธ์>}}</tool_call>
์ฌ์ฉ ๊ฐ๋ฅํ ํด ๋ชฉ๋ก:
<tools>
{tools_json}
</tools>
{tools_json} ์๋ฆฌ์ ๋ค์ด๊ฐ๋ ํด ์คํค๋ง ํ์ (OpenAI function ์คํค๋ง์ ๋์ผํ parameters ๊ตฌ์กฐ):
{"name": "get_weather", "description": "์ง์ญ์ ํ์ฌ ๋ ์จ๋ฅผ ์กฐํํ๋ค", "parameters": {"type": "object", "properties": {"location": {"type": "string", "description": "๋์ ์ด๋ฆ"}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}}, "required": ["location"]}}
{"name": "reply", "description": "ํด์ด ํ์ ์๋ ์ง๋ฌธ์ ์ง์ ๋ตํ๋ค", "parameters": {"type": "object", "properties": {"text": {"type": "string"}}, "required": ["text"]}}
{"name": "escalate", "description": "์ฒ๋ฆฌํ ์ ์๋ ์์ฒญ์ ์์๋ก ๋๊ธด๋ค", "parameters": {"type": "object", "properties": {"query": {"type": "string"}}, "required": ["query"]}}
reply์ escalate๋ ํญ์ ๋ชฉ๋ก์ ํฌํจํด์ผ ํฉ๋๋ค. ๋๋จธ์ง ํด์ ์์ ๋กญ๊ฒ ๊ต์ฒดํ ์ ์๊ณ , ํ์ต ๋ ๋ณธ ์ ์๋ ํด์
๋ ๋์ํฉ๋๋ค(zero-shot). ํด ์ค๋ช
์ ์์ดยทํ๊ตญ์ด ๋ชจ๋ ์ง์ํฉ๋๋ค.
๊ธฐ๋ณธ ํด 3์ข (๋ผ์ฐํฐ์ ํ๋กํ ์ฝ)
์ด ๋ผ์ฐํฐ๋ "๋ชจ๋ ์๋ต์ด ํด์ฝ"์ด๋ผ๋ ๊ท์ฝ ์์์ ๋์ํ๋ฉฐ, ๋๋ฉ์ธ ํด ์ธ์ ๋ค์ ๊ธฐ๋ณธ ํด๋ก ๋ชจ๋ ๋ฐํ๋ฅผ ์ปค๋ฒํฉ๋๋ค:
| ๊ธฐ๋ณธ ํด | ์ธ์ ๋์ค๋ | ์ธ์ | ์๋ฒ๊ฐ ํ ์ผ |
|---|---|---|---|
reply |
ํด์ด ํ์ ์๋ ์ธ์ฌยท์ก๋ดยท์์ ์ง๋ฌธ ("๊ณ ๋ง์!", "์๋ ") | text: ์งง์ ํ๊ตญ์ด ๋ต๋ณ |
text๋ฅผ ๊ทธ๋๋ก ์ฌ์ฉ์์๊ฒ |
escalate |
๋ชฉ๋ก์ ํด๋ก ์ฒ๋ฆฌ ๋ถ๊ฐํ ์์ฒญ โ ๋ค๋ฅธ ๋๋ฉ์ธ, ๋ณตํฉ ์์ , ํ์ ์ธ์๋ฅผ ์ ์ ์์, ๊ธด ๊ธ์ฐ๊ธฐ ๋ฑ | query: ์ฌ์ฉ์ ์์ฒญ ์๋ฌธ ๊ทธ๋๋ก |
์์ ๋ชจ๋ธยท๋ด๋น์์๊ฒ ์ ๋ฌ |
recall |
๋ฐํ๊ฐ ๊ณผ๊ฑฐ ๋ํ์ ๊ฐ์ ๊ฐ๋ฆฌํฌ ๋ ("๊ทธ๊ฑฐ ์ฌ๊ณ ์ผ๋ง ๋จ์์ด?", "์๊น ๊ทธ ๋์ ๋ด์ผ ๋ ์จ๋?") | needs: ์ฐพ์์ผ ํ ๊ฐ์ ํ๊ตญ์ด ์ค๋ช
|
recallResolver-1B๋ก ํ์คํ ๋ฆฌ์์ ๊ฐ์ ์ฐพ์ ๋ฐํ๋ฅผ ๋ณด๊ฐ ํ ๋ผ์ฐํฐ ์ฌํธ์ถ |
์ค๊ณ ์๋: ๋ผ์ฐํฐ๋ ์์ ํ ์คํธ๋ฅผ ์์ฑํ์ง ์์ผ๋ฏ๋ก "๋๋ตํ๋ค/๋๊ธด๋ค/๊ณผ๊ฑฐ๋ฅผ ์ฐพ๋๋ค"๊น์ง ์ ๋ถ ํด์ฝ๋ก ํํ๋ฉ๋๋ค. ๋๋ถ์ ๋ค์ด์คํธ๋ฆผ์ ๋ถ๊ธฐ ์ฒ๋ฆฌ๋ง ํ๋ฉด ๋๊ณ , ์ถ๋ ฅ ๋ฌธ๋ฒ ๊ฐ์ (constrained decoding)๋ฅผ ์ ์ฒด ์๋ต์ ์ ์ฉํ ์ ์์ต๋๋ค.
reply ์ด์ฉ ํ: reply.text๋ฅผ ๊ทธ๋๋ก ์ธ ์๋ ์์ง๋ง(์ง์ฐ ์ต์), reply๋ฅผ "๋ํ ํธ๋" ์ ํธ๋ก๋ง ์ฐ๊ณ ๋ฐํ๋ฅผ ๋ณ๋ ์ฑ ๋ชจ๋ธ๋ก ๋๊ธฐ๋ ๊ตฌ์ฑ๋ ๋ถ๊ธฐ ํ ์ค์ด๋ฉด ๋ฉ๋๋ค. ํ๋ฅด์๋ ์๋ ๋ณธ๊ฒฉ ์ฑ๋ด๊ณผ ๋ถ์ผ ๋ ์ ์ฉํฉ๋๋ค:
call = route(utterance)
if call["name"] == "reply":
return chat_model(persona, utterance) # text๋ ๋ฒ๋ฆฌ๊ณ ํฐ ๋ชจ๋ธ์ด ๋ํ
recall ์ฐ๋ ํ (recallResolver-1B์์ end-to-end ๊ฒ์ฆ์์ ํ์ ํ ์๋ฒ ๊ท์น): โ ๋ฆฌ์กธ๋ฒ๊ฐ ๊ฐ์ ์ฐพ์ผ๋ฉด ๊ทธ ๊ฐ์ ์ถ์ฒ ์ธ์ ์คํค๋ง ์ค๋ช ์ ๋ผ๋ฒจ๋ก ๋ฐํ๋ฅผ ๋ณด๊ฐํด ์ฌ๋ผ์ฐํ โ
"{๋ฐํ} ({๋ผ๋ฒจ}: {๊ฐ})"โก๋ผ์ฐํฐ์ needs๋ก ๋ฆฌ์กธ๋ฒ๊ฐ ๋ชป ์ฐพ์ผ๋ฉด ํด๋น ๊ธฐ๋ก์ ์ธ์ ์ค๋ช ๋ค๋ก ์ฌ์ง์ โขneeds์ ๊ฐ ๋ผ๋ฒจ์ ์ ํ์ด ์ ๋ง์ผ๋ฉด ์๋ ์ฃผ์ ๋์ ์ฌ์ฉ์์๊ฒ ํ์ธ. ์ฌํธ์ถ์ด ๋ recall์ด๋ฉด escalate(๋ฃจํ ๊ฐ๋). recall์ ์ฐ์ง ์์ผ๋ ค๋ฉด ํด ๋ชฉ๋ก์์ ๋นผ๋ฉด ๋ฉ๋๋ค โ ๊ทธ๋ฌ๋ฉด ๊ณผ๊ฑฐ ์ฐธ์กฐ ๋ฐํ๋ escalate๋ก ๋์ต๋๋ค.
๋์ ์:
user: ์ฐ์ฐ ์ฑ๊ฒจ์ผ ํ๋, ๋์ ์ธ๋ฐ
assistant: <tool_call>{"name": "get_weather", "arguments": {"location": "๋์ "}}</tool_call>
user: ๊ณ ๋ง์!
assistant: <tool_call>{"name": "reply", "arguments": {"text": "๋ณ๋ง์์์, ๋ ํ์ํ๋ฉด ๋ถ๋ฌ ์ฃผ์ธ์!"}}</tool_call>
user: ์ด ์ฌ์ง์์ ๊ธ์ ์ข ์ฝ์ด์ค
assistant: <tool_call>{"name": "escalate", "arguments": {"query": "์ด ์ฌ์ง์์ ๊ธ์ ์ข ์ฝ์ด์ค"}}</tool_call>
GGUF (llama.cpp)
์ด๋ํฐ๋ฅผ ๋จธ์งํด ์์ํํ ๋จ์ผ ํ์ผ๋ ์ ๊ณตํฉ๋๋ค. ์์ค์ ์์ฒด ํ๊ฐ์ ์ค์ธก์น์ ๋๋ค:
| ํ์ผ | ํฌ๊ธฐ | ํ๊ตญ์ด ํ๊ฐ์ (decision/tool/args) | ๊ถ์ฅ ์ฉ๋ |
|---|---|---|---|
toolRouter-1B-Q8_0.gguf |
1.3GB | 96.5 / 95.0 / 83.4 (์๋ณธ ๋๋น ~-0.8%p) | ํ์ง ์ฐ์ |
toolRouter-1B-Q4_K_M.gguf |
0.8GB | 95.5 / 94.2 / 83.4 (~-1.6%p) | ์ต์ VRAM |
llama-server -m toolRouter-1B-Q8_0.gguf -ngl 99 -c 2048
์ฑํ
ํ
ํ๋ฆฟ์ด GGUF์ ํฌํจ๋ผ ์์ต๋๋ค. greedy(temperature 0) ๊ถ์ฅ, ์ถ๋ ฅ์ <tool_call>...</tool_call> ํ์ ๊ทธ๋๋ก์
๋๋ค. recall ํ์ ์ ์์ํ์ ๋ค์ ๋ฏผ๊ฐํ๋(๊ฒฝ๊ณ ์ผ์ด์ค -4~5%p) recall์ ์ฐ๋ ํ์ดํ๋ผ์ธ์ด๋ฉด Q8์ ๊ถํฉ๋๋ค.
์ ํ
- ๋จ์ผ ํด ๋ผ์ฐํ ์ ์ฉ์ ๋๋ค (๋ฉํฐํด ๋ฌธ๋งฅ ์ฐธ์กฐ๋ ๋ฏธ์ง์).
- ์ธ๊ณ์ง์ ์ถ๋ก ์ด ํ์ํ ์ธ์ ๋ณด์ (์: ๋์๋ช โํตํ์ฝ๋), ์๋ ๋ ์ง ํด์์ ๋ค์ด์คํธ๋ฆผ์์ ์ฒ๋ฆฌํ์ธ์. ์๋ ๋ ์ง๋ ์์คํ ํ๋กฌํํธ์ ํ์ฌ ๋ ์ง๋ฅผ ์ ๊ณตํ๋ฉด ๊ฐ์ ๋ฉ๋๋ค.
- base ๋ชจ๋ธ์ Llama 3.2 Community License๋ฅผ ๋ฐ๋ฆ ๋๋ค.
- Downloads last month
- -
4-bit
8-bit
Model tree for kimjg/toolRouter-1B
Base model
meta-llama/Llama-3.2-1B-Instruct