Instructions to use Smriti10raj/support-slm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Smriti10raj/support-slm with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct") model = PeftModel.from_pretrained(base_model, "Smriti10raj/support-slm") - Notebooks
- Google Colab
- Kaggle
You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
This model reads computer-support messages for one application (TicketAI) and is published as evidence for a method, not as a general-purpose assistant. Tell us who you are and what you intend to use it for.
Log in or Sign Up to review the conditions and access this model content.
support-slm (0.5B): v1, and v2 as an off-topic judge
A LoRA adapter for Qwen2.5-0.5B-Instruct that reads one computer-support message (typed, voice-typed, messy, or Hinglish) and returns one compact JSON object describing it:
{"intent":"technical_support","vendor":"Dell","family":"Latitude","model":"5420","issue":"BSOD","symptoms":["after waking from sleep"],"codes":["DRIVER_POWER_STATE_FAILURE"],"tried":["driver update"],"action":"alternative_solution","sentiment":"neutral","priority":"medium"}
It is the "understand" step of TicketAI, a local support assistant. It never answers the technical question itself: the application searches real vendor articles for that, and plain code decides what to do next. It is published as evidence for a method, not as a model to reuse: it learned one label vocabulary (below) inside one exact system prompt.
How to use it
The system prompt must be the one it was trained with, byte for byte: system_prompt.txt.
Decode greedily.
import torch
from huggingface_hub import hf_hub_download
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = "Qwen/Qwen2.5-0.5B-Instruct"
rev = "7ae557604adf67be50417f59c2c2f167def9a775"
tok = AutoTokenizer.from_pretrained(base, revision=rev)
model = PeftModel.from_pretrained(AutoModelForCausalLM.from_pretrained(base, revision=rev, torch_dtype=torch.float32),
"Smriti10raj/support-slm", revision="v1").eval()
system = open(hf_hub_download("Smriti10raj/support-slm", "system_prompt.txt", revision="v1")).read().strip()
msgs = [{"role": "system", "content": system},
{"role": "user", "content": "my dell latitude 5420 blue screens after sleep, already updated drivers"}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt")
out = model.generate(ids, max_new_tokens=160, do_sample=False)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))
A faster llama.cpp build (about 4× faster on CPU, same or better scores) is in
Smriti10raj/support-slm-GGUF.
Output vocabulary
intent: technical_support | greeting | out_of_scope. issue, action, symptoms and tried come from
fixed lists written in the system prompt (14 issues such as BSOD, No boot, Not charging, Wi-Fi; 21
actions such as suggest_solution, alternative_solution, escalate, draft_support_email). Keys the message doesn't support are
left out: the model is trained to omit, not guess.
Results
Three test sets, none used in training. Training rows that were exact, normalised, fuzzy (≥ 0.90) or semantic (cosine ≥ 0.95) near-duplicates of any test row were removed before training.
- Held-out (420 messages): generated like the training data but with phrasings and device models that never appear in training.
- Hand-written (45 messages): long stories, experts, voice-typed text, Hinglish, multi-part requests and hard negatives.
- Refusal test (60 messages): 40 off-topic messages, then 20 real support messages that must not be refused.
Held-out test, the model alone
| support-slm v1 (PyTorch adapter) | Q8_0 GGUF + LoRA (llama.cpp) | rules only | untuned Qwen2.5-0.5B | |
|---|---|---|---|---|
| Valid JSON | 1.0 | 1.0 | 1.0 | 0.964 |
| Whole record right (every field) | 0.481 | 0.555 | 0.162 | 0.0 |
| Intent | 1.0 | 1.0 | 0.836 | 0.89 |
| Issue | 0.717 | 0.756 | 0.349 | 0.056 |
| Requested action | 0.906 | 0.923 | 0.352 | 0.44 |
| Vendor | 0.955 | 0.963 | 0.96 | 0.414 |
| Product model | 0.976 | 0.976 | 0.876 | 0.122 |
| Error codes (F1) | 0.932 | 0.965 | 0.938 | 0.0 |
| Symptoms (F1) | 0.454 | 0.532 | 0.021 | 0.009 |
| Already-tried steps (F1) | 0.695 | 0.682 | 0.0 | 0.0 |
| Invented a vendor/model/code (rate) | 0.003 | 0.0 | 0.007 | 0.56 |
| Seconds per message (CPU) | 6.649 | 1.689 | 0.003 | 15.324 |
Hand-written test and refusal test
| support-slm v1 (llama.cpp) | rules only | |
|---|---|---|
| Hand-written: whole record right | 0.4 | 0.267 |
| Hand-written: off-topic or greeting handled correctly | 4/5 | 4/5 |
| Refusal test: off-topic refused | 24/40 | 26/40 |
| Refusal test: real support wrongly refused | 0/20 | 3/20 |
Inside TicketAI
The application combines the model with small rules (exact device names and explicit requests like "escalate" come from the rules; labels outside the vocabulary are mapped or dropped). Measured through that real code: whole record right 0.536 held-out and 0.444 hand-written. On the refusal test it refuses 21/40 off-topic messages and wrongly refuses 0/20 real ones.
v2: better at turning away off-topic questions (used as a judge next to v1)
v2 is the same recipe (Qwen2.5-0.5B-Instruct, LoRA r16, alpha 32, 2 epochs, laptop CPU, 19.4 hours)
trained on support-v2-synth: 3211 training rows that add general-tech questions, buying advice
and non-computer requests labelled out_of_scope, plus counterweight rows so real support is not refused.
Final train loss 0.0255, eval loss 0.0073. It is in the v2/ folder of this repo
(tag v2); the GGUF LoRA is support-slm-v2-lora-F16.gguf in Smriti10raj/support-slm-GGUF.
model = PeftModel.from_pretrained(base_model, "Smriti10raj/support-slm", subfolder="v2", revision="v2")
The system prompt is identical to v1's. v2 was evaluated only as the llama.cpp build (Q8_0 base + F16 LoRA).
The model alone (llama.cpp)
| v1 | v2 | |
|---|---|---|
| Held-out (420): whole record right | 0.555 | 0.502 |
| Hand-written (45): whole record right | 0.4 | 0.533 |
| Refusal test: off-topic refused | 24/40 | 40/40 |
| Refusal test: real support wrongly refused | 0/20 | 1/20 |
| Held-out: priority right | 0.94 | 0.855 |
Inside TicketAI
| v1 + rules | v2 + rules | v1, with v2 as off-topic judge | |
|---|---|---|---|
| Held-out: whole record right | 53.6% | 48.1% | 53.3% |
| Hand-written: whole record right | 44.4% | 51.1% | 46.7% |
| Off-topic turned away | 21/40 | 37/40 | 33/40 |
| Real support wrongly refused | 0/20 | 1/20 | 0/20 |
Why v2 did not replace v1. A rule fixed before v2 was trained said v2 replaces v1 only if it is no worse on the
old tests. It is worse on the held-out set, mostly because it rates urgent messages medium instead of high
priority. TicketAI therefore keeps v1 for every field and runs v2 in parallel as a judge: v2's answer is used only
when v2 says out_of_scope, v1 does not, and the message names no vendor, product line, known issue or error code.
That last guard was added after a live check in which v2 alone turned away a voice-typed Dell blue-screen message;
the judge figures above were measured after it was added, on the models' saved outputs through TicketAI's code.
Limits (read before using)
- It refuses too few off-topic questions. General tech questions ("SSD vs HDD, which is better?", buying advice that names Dell or HP) are often read as support requests. v2 (below) turns away far more of them and is used in TicketAI as a judge next to v1.
- Symptoms are the weakest field (F1 0.454–0.532 held-out); already-tried steps reach F1 0.682. Treat both as hints.
- It only knows the vocabulary above: Dell, HP, Lenovo, Microsoft (Surface), Windows laptops and desktops, English and Hinglish. Other vendors, languages or device types are out of scope.
- Synthetic training data. Labels are correct by construction, but real users will write things the generator never produced. The hand-written test is small (45 messages).
- It sees one message at a time. Conversation state lives in the application, not in the model.
Training (v1)
- Base:
Qwen/Qwen2.5-0.5B-Instruct@7ae557604adf. LoRA r16, alpha 32, loss on the answer only, 2 epochs, on a laptop CPU (about 15 hours). - Data:
support-v1-synth: 3181 training and 338 validation rows, generated from slot templates with split phrasing banks (validation uses whole phrasing groups never seen in training). Train sha256123ddbc0374758f6…. - Final train loss 0.0239, eval loss 0.0731.
adapter_model.safetensorssha256cd291a325e6a5e5f….
- Downloads last month
- 1