Instructions to use phonology024/localllm-task-router with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use phonology024/localllm-task-router with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf phonology024/localllm-task-router:Q8_0 # Run inference directly in the terminal: llama cli -hf phonology024/localllm-task-router:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf phonology024/localllm-task-router:Q8_0 # Run inference directly in the terminal: llama cli -hf phonology024/localllm-task-router:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf phonology024/localllm-task-router:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf phonology024/localllm-task-router:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf phonology024/localllm-task-router:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf phonology024/localllm-task-router:Q8_0
Use Docker
docker model run hf.co/phonology024/localllm-task-router:Q8_0
- LM Studio
- Jan
- Ollama
How to use phonology024/localllm-task-router with Ollama:
ollama run hf.co/phonology024/localllm-task-router:Q8_0
- Unsloth Desktop
- Docker Model Runner
How to use phonology024/localllm-task-router with Docker Model Runner:
docker model run hf.co/phonology024/localllm-task-router:Q8_0
- Lemonade
How to use phonology024/localllm-task-router with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull phonology024/localllm-task-router:Q8_0
Run and chat with the model
lemonade run user.localllm-task-router-Q8_0
List all available models
lemonade list
- Atomic Chat
localllm task router (multilingual-e5-small Q8_0 + classifier head)
The task classifier used by make-localllm-easier (localllm serve --models auto). It labels each chat message as general / math / code / translate in any language, and the router then picks the local model with the best measured score for that task.
| file | what |
|---|---|
multilingual-e5-small-Q8_0.gguf (126 MB) |
intfloat/multilingual-e5-small (MIT), converted with llama.cpp convert_hf_to_gguf.py and quantized to Q8_0. Embeddings match the PyTorch model at cosine 0.992–0.9998. |
router_head.json (31 KB) |
4 × 384 logistic-regression head trained on those embeddings (L2-normalised, mean pooling, input prefixed query: , first 450 characters). |
Use with llama.cpp: llama-server -m multilingual-e5-small-Q8_0.gguf --embedding --pooling mean -ngl 0, then compute softmax(W·x/|x| + b) with the head. localllm/taskclf.py does this for you.
Measured (shared test: 144 messages in 15 languages):
| accuracy | median latency (CPU) | size | |
|---|---|---|---|
| this router (keyword rules first) | 98.6% | ~5–8 ms | 126 MB |
| embedding only | 97.9% | ||
| Laya multilingual (zero-shot) | 77.8% | 103 ms | 614 MB |
Caveat: the shared test set and part of the training data are AI-written, so a human-written test set is still being collected. The training data, scripts and test set are in tools/router/.
The file and the head are a matched pair: retrain the head if you re-convert the GGUF.
- Downloads last month
- 23
8-bit
Model tree for phonology024/localllm-task-router
Base model
intfloat/multilingual-e5-small