localllm task router (multilingual-e5-small Q8_0 + classifier head)

The task classifier used by make-localllm-easier (localllm serve --models auto). It labels each chat message as general / math / code / translate in any language, and the router then picks the local model with the best measured score for that task.

file what
multilingual-e5-small-Q8_0.gguf (126 MB) intfloat/multilingual-e5-small (MIT), converted with llama.cpp convert_hf_to_gguf.py and quantized to Q8_0. Embeddings match the PyTorch model at cosine 0.992–0.9998.
router_head.json (31 KB) 4 × 384 logistic-regression head trained on those embeddings (L2-normalised, mean pooling, input prefixed query: , first 450 characters).

Use with llama.cpp: llama-server -m multilingual-e5-small-Q8_0.gguf --embedding --pooling mean -ngl 0, then compute softmax(W·x/|x| + b) with the head. localllm/taskclf.py does this for you.

Measured (shared test: 144 messages in 15 languages):

accuracy median latency (CPU) size
this router (keyword rules first) 98.6% ~5–8 ms 126 MB
embedding only 97.9%
Laya multilingual (zero-shot) 77.8% 103 ms 614 MB

Caveat: the shared test set and part of the training data are AI-written, so a human-written test set is still being collected. The training data, scripts and test set are in tools/router/.

The file and the head are a matched pair: retrain the head if you re-convert the GGUF.

Downloads last month
23
GGUF
Model size
0.1B params
Architecture
bert
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for phonology024/localllm-task-router

Quantized
(290)
this model