TinyJudge (v2.1)
A 150M-parameter fine-tuned judge for AI-agent behaviour: a fast, calibrated alternative to using a large LLM as the judge for yes/no evaluation questions. Code, data pipeline and API: github.com/Charansripadi/tinyjudge.
It answers three questions about an agent's tool-using conversation:
| Task | Question |
|---|---|
policy |
Did the agent's actions follow the rules? |
tool_choice |
Is the proposed next tool call appropriate? |
grounded |
Is every fact in the agent's reply supported by its tool results? |
Outputs are calibrated probabilities (Platt scaling per task, fitted on held-out real agent traffic), and each task has a cascade margin: inside it, the model is not confident enough and the case should go to an LLM judge.
Results on real agent conversations
208 judgements from real LLM-agent conversations (SkillMiner's support agent), never used for training or for fitting calibration:
| Task | n | TinyJudge (calibrated) | LLM judge (Qwen 3.8 27B) | Cascade | Escalated to LLM |
|---|---|---|---|---|---|
| policy | 38 | 97.4% | 71.1% | 92.1% | 16% |
| tool_choice | 85 | 88.2% | 87.1% | 92.9% | 8% |
| grounded | 85 | 83.5% | 80.0% | 85.9% | 71% |
| all | 208 | 88.0% | 81.2% | 89.9% | 35% |
Calibration on the same real test half (ECE = expected calibration error, lower is better): grounded 0.201 โ 0.101, tool_choice 0.138 โ 0.073, policy 0.049 โ 0.078.
Usage
from tinyjudge.predictor import Judge
from tinyjudge.tasks import Action, grounded_text
judge = Judge("CharanSripadi/tinyjudge", revision="v2.1")
text = grounded_text([Action("track_shipment", {"order_id": "A1002"},
{"status": "ok", "shipping_status": "shipped", "eta": "2 days"})],
"Your order A1002 has shipped and should arrive in 2 days.")
print(judge.judge([("grounded", text)]))
Inputs must be formatted with tinyjudge.tasks (policy_text, tool_choice_text, grounded_text),
the same functions used in training. Or run the API: docker run -p 7860:7860 ghcr.io/charansripadi/tinyjudge.
Training
- Base:
answerdotai/ModernBERT-base, sequence classification (no / yes), one model for all three tasks. - Data v2: 2,806 training examples, generated from a simulated shop with rule-derived (exact) labels: correct and mistaken agent trajectories, template and LLM-written replies, one-fact corruptions.
- 3 epochs, lr 3e-5, batch 16, fp16, max length 512. About 4 minutes on a free Colab T4.
- Tracked with MLflow; the dataset manifest's SHA-256 is logged with every run.
Limitations
- One domain. Trained on a demo online-shop support agent with six tools. Other domains need new data.
- Small real test set. 38 to 85 examples per task; differences of a few points are within noise.
- Label assumptions. Real
groundedpositives assume the agent's own reply matched its tool results; at least one such reply contained an invented timeline, so some "errors" are the model being right. - LLM baseline. Qwen 3.8 27B without reasoning. A stronger or reasoning model would likely do better on
policy. - Cascade. Escalating
policyhurt accuracy here, because the LLM is worse than TinyJudge on that task.
- Downloads last month
- -
Model tree for CharanSripadi/tinyjudge
Base model
answerdotai/ModernBERT-base