SITOR - Intelligent Triage and Routing Engine (RoBERTa)
Model Description
SITOR (Sistema Inteligente de Tipificaci贸n, Orquestaci贸n y Resoluci贸n de Incidencias) is a fine-tuned roberta-base (125M) model designed for automated BPO (Business Process Outsourcing) and IT Helpdesk ticket routing.
Unlike standard classification models that attempt to predict blindly, SITOR is calibrated to act as an algorithmic firewall. It predicts over 56 highly overlapping operational queues but is designed to only automate routing when the mathematical certainty (Softmax output calibrated via L-BFGS) exceeds a rigorous 0.85 threshold.
- Developed by: Pedro Andreu Torres (Master in Data Science & AI Thesis)
- Model Type: Transformer-based sequence classifier
- Language(s): English (Trained on English telecom support corpus)
- Base Model:
roberta-base
Intended Use & Business Logic
This model is intended for enterprise-level ticket orchestration.
Due to the asymmetric cost of misrouting a ticket in a Tier-1 BPO environment, the model does not operate on a simple argmax decision.
- High Confidence (>= 0.85): The ticket is strictly automated and bypassed to the designated technical queue.
- Low Confidence (< 0.85): The model acknowledges ambiguity and flags the ticket for manual human review.
Training Data
To ensure corporate confidentiality and academic rigor, the model was trained on a public Telecommunications Customer Support dataset sourced from Kaggle. The raw data underwent severe deduplication and asymmetric mapping to simulate a complex 56-queue BPO ecosystem.
Evaluation Results
Validation was performed using a strict StratifiedGroupKFold strategy to prevent Data Leakage on a blind hold-out set of 3,080 tickets.
- Base Algorithmic Accuracy: ~59% (Top-1 Accuracy on 56 highly overlapping classes).
- Automated Volume (0.85 Threshold): 38.5% of the total influx.
- Operational Precision (AI): On the automated volume, the model achieves 97.3% precision (only 2.7% residual error).
- In-Vitro ROI: The hybrid approach yields a theoretical 31.8% reduction in global operational errors compared to a 100% human-operated baseline. (Note: Requires Shadow Mode validation in a live production environment).
How to use
You can load this model and its tokenizer directly using the transformers library:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
# Load the model and tokenizer from the Hugging Face Hub
tokenizer = AutoTokenizer.from_pretrained("Andregon79/sitor")
model = AutoModelForSequenceClassification.from_pretrained("Andregon79/sitor")
# Example ticket
ticket_text = "My home wifi keeps dropping every 10 minutes since yesterday."
inputs = tokenizer(ticket_text, return_tensors="pt", truncation=True, max_length=128)
with torch.no_grad():
logits = model(**inputs).logits
# Apply softmax to get confidence scores
probabilities = torch.softmax(logits, dim=1)
confidence, predicted_class = torch.max(probabilities, dim=1)
print(f"Confidence: {confidence.item():.4f}")
- Downloads last month
- 18