nepkev

Kev 0.8B fine-tuned to route Nepali customer-support messages for three demo services: an online shop, a digital wallet and a bank. Messages can be Devanagari or Romanized Nepali, often mixed with English. For each message the model answers two questions:

  • q1, issue: probabilities over the service's categories (other for clear requests outside the list);
  • q2, clarification: the probability that the agent must ask a follow-up question first.

Instructions and category descriptions are English; only the message is Nepali. The categories are a demo taxonomy, not any real institution's. The model only routes.

Code, taxonomy and training pipeline: github.com/sm079/nepkev.

Results

On a held-out test set of 889 generated messages (95% CI from a scenario-group bootstrap):

issue accuracy clarification F1
keyword rules 0.619 0.306
Kev-0.8B, untouched 0.425 0.182
nepkev 0.966 [0.953, 0.980] 0.874 [0.797, 0.933]

Fine-tuning adds +0.541 issue accuracy [+0.496, +0.588] over the untouched model on the same messages. By part of the test set (issue accuracy):

untouched fine-tuned
clean Devanagari 0.301 0.985
Romanized 0.420 0.952
app-review style 0.557 0.962

Issue ECE 0.021 after temperature calibration. No category is over-predicted (other at 1.03x its true share, precision 0.96).

Training

About 7,400 labelled messages, balanced across categories: casual Devanagari, its Romanized twins, the same situations in app-review style, and a small set of labelled app reviews. LoRA fine-tuning from jaredpalmer/kev-0.8b@9a45d25 with Kev 5e42a7a, 2 epochs (the epoch with the lowest dev NLL is kept), temperature fitted on a calibration split.

Use

A Kev checkpoint: adapter_model.safetensors is the LoRA adapter on Qwen/Qwen3.5-0.8B-Base@dc7cdfe, head.pt holds the pointer head and the fitted temperature. With Kev at 5e42a7a:

python -m kev.serve --run sm079/nepkev        # POST /v1/systemone

The request format (question wording and category descriptions per service) is in the code's configs/taxonomy/demo-1.yaml.

In the browser

onnx/ holds the same model for ONNX Runtime Web (WebGPU, or CPU through WebAssembly). The adapter is merged into the backbone, which is exported without its language-model head so that it outputs hidden states; head.bin is the pointer head (float32) and manifest.json lists the files and graph inputs. Each question runs as one token row, <|fim_prefix|> message <|fim_middle|> instruction (<|box_start|> option <|box_end|>)... <|fim_suffix|>, and the head scores every option from the hidden states at <|fim_suffix|> and at each <|box_end|>.

folder download answers that differ from full precision (240 test questions) issue accuracy (108 test questions; full precision 0.972)
onnx/int8 707 MB 3 0.963
onnx/int8-fp16emb 1,055 MB 1 0.972

int8 also stores the embedding table in 4-bit blocks. 4-bit weights throughout cost about 5 points of issue accuracy (0.917), so they are not published. A static demo page that runs these files is in the code's demo/ folder.

Downloads last month
2
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sm079/nepkev

Adapter
(42)
this model