Instructions to use gilangrp/support_ticket_llm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use gilangrp/support_ticket_llm with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf gilangrp/support_ticket_llm # Run inference directly in the terminal: llama cli -hf gilangrp/support_ticket_llm
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf gilangrp/support_ticket_llm # Run inference directly in the terminal: llama cli -hf gilangrp/support_ticket_llm
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf gilangrp/support_ticket_llm # Run inference directly in the terminal: ./llama-cli -hf gilangrp/support_ticket_llm
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf gilangrp/support_ticket_llm # Run inference directly in the terminal: ./build/bin/llama-cli -hf gilangrp/support_ticket_llm
Use Docker
docker model run hf.co/gilangrp/support_ticket_llm
- LM Studio
- Jan
- Ollama
How to use gilangrp/support_ticket_llm with Ollama:
ollama run hf.co/gilangrp/support_ticket_llm
- Unsloth Desktop
- Docker Model Runner
How to use gilangrp/support_ticket_llm with Docker Model Runner:
docker model run hf.co/gilangrp/support_ticket_llm
- Lemonade
How to use gilangrp/support_ticket_llm with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull gilangrp/support_ticket_llm
Run and chat with the model
lemonade run user.support_ticket_llm-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Fine-Tuning Phi-3-mini for Support Ticket Extraction
Fine-tunes an LLM with Unsloth to turn free-text customer complaints/emails
into structured JSON: product, category, urgency, sentiment.
Model
- Base model:
unsloth/Phi-3-mini-4k-instruct-bnb-4bit - Microsoft's Phi-3-mini (3.8B params, instruction-tuned), re-packaged by Unsloth in 4-bit quantized form for lightweight fine-tuning (e.g. free Colab T4 GPU).
Files
| File | Description |
|---|---|
Fine-Tuning_Unsloth_SupportTicket.ipynb |
Main notebook β load model, dataset, LoRA, training, inference, GGUF export |
support_ticket_data.json |
Training data (12 prompt β JSON examples) |
inferece_test.txt |
Extra test cases (brands not in training data) to check generalization |
Modelfile |
Ollama config auto-generated by Unsloth |
Workflow (Colab)
- Runtime β Change runtime type β Python 3 + T4 GPU
!pip install unsloth- Load model with
FastLanguageModel.from_pretrained(...) - Format dataset into chat-template text (
tokenizer.apply_chat_template) - Apply LoRA (
r=64, attention + MLP layers) - Train with
SFTTrainer(max_steps=60, batch size 2, grad accumulation 4) - Quick test with
model.generate() - Export to GGUF:
model.save_pretrained_gguf(..., quantization_method="q4_k_m")
Running Locally (Ollama)
brew install ollama
ollama create support-ticket-phi3 -f Modelfile
ollama run support-ticket-phi3
Test with multiple cases:
while IFS= read -r line; do
echo "--- Input: $line ---"
ollama run support-ticket-phi3 "$line"
echo ""
done < inferece_test.txt
Result test case:
ollama run support-ticket-phi3
>>> I ordered a MALM bed frame last week and one of the side panels arrived with a big scratch. Not a huge deal but I'd like a replacement panel sent over.
{"category": "product defect", "product": "MALM bed frame", "sentiment": "neutral", "urgency": "low"}
ollama run support-ticket-phi3
>>> Loved how fast your support team replied when I asked about my Adidas Ultraboost order, really appreciated the quick and friendly help.
{"category": "feedback", "product": "Adidas Ultraboost", "sentiment": "positive", "urgency": "low"}
ollama run support-ticket-phi3
>>> The Samsung Galaxy Buds I bought stopped connecting to my phone after only a few days. Pretty frustrating since I use them daily for calls.
{"category": "product defect", "product": "Samsung Galaxy Buds", "sentiment": "negative", "urgency": "medium"}
Known Limitations
- Tiny dataset (12 examples) β fine for demo purposes, but not enough for real generalization. Scale up to hundreds of varied examples for production.
- Model loses general-purpose ability after fine-tuning β since every
training example follows "free text β JSON", the model always replies in
JSON, even for unrelated questions (e.g. "who is Isaac Newton?" still
returns a JSON blob). This isn't a bug β it's expected from small,
single-pattern fine-tuning. Fix by mixing in:
- Explicit instructions per prompt ("Extract this ticket into JSON: ...")
- Regular conversational examples alongside extraction examples
- A consistent system prompt at both training and inference time
- Default
temperature = 1.5in the auto-generatedModelfileis too high for structured extraction β lower it to0.1β0.3for consistent JSON.
Compatibility Notes
- Training requires an NVIDIA GPU (CUDA) β won't run on Mac/Apple Silicon,
since
unslothandbitsandbytesdepend on Triton/CUDA. Use Colab's free T4 GPU or another CUDA cloud. - Inference on the final
.gguffile is lightweight and runs fine on CPU/Apple Silicon via Ollama or llama.cpp β no GPU needed.
- Downloads last month
- 14
Hardware compatibility
Log In to add your hardware
We're not able to determine the quantization variants.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support