TinyJudge (v2.1)

A 150M-parameter fine-tuned judge for AI-agent behaviour: a fast, calibrated alternative to using a large LLM as the judge for yes/no evaluation questions. Code, data pipeline and API: github.com/Charansripadi/tinyjudge.

It answers three questions about an agent's tool-using conversation:

Task Question
policy Did the agent's actions follow the rules?
tool_choice Is the proposed next tool call appropriate?
grounded Is every fact in the agent's reply supported by its tool results?

Outputs are calibrated probabilities (Platt scaling per task, fitted on held-out real agent traffic), and each task has a cascade margin: inside it, the model is not confident enough and the case should go to an LLM judge.

Results on real agent conversations

208 judgements from real LLM-agent conversations (SkillMiner's support agent), never used for training or for fitting calibration:

Task n TinyJudge (calibrated) LLM judge (Qwen 3.8 27B) Cascade Escalated to LLM
policy 38 97.4% 71.1% 92.1% 16%
tool_choice 85 88.2% 87.1% 92.9% 8%
grounded 85 83.5% 80.0% 85.9% 71%
all 208 88.0% 81.2% 89.9% 35%

Calibration on the same real test half (ECE = expected calibration error, lower is better): grounded 0.201 โ†’ 0.101, tool_choice 0.138 โ†’ 0.073, policy 0.049 โ†’ 0.078.

Usage

from tinyjudge.predictor import Judge
from tinyjudge.tasks import Action, grounded_text

judge = Judge("CharanSripadi/tinyjudge", revision="v2.1")
text = grounded_text([Action("track_shipment", {"order_id": "A1002"},
                             {"status": "ok", "shipping_status": "shipped", "eta": "2 days"})],
                     "Your order A1002 has shipped and should arrive in 2 days.")
print(judge.judge([("grounded", text)]))

Inputs must be formatted with tinyjudge.tasks (policy_text, tool_choice_text, grounded_text), the same functions used in training. Or run the API: docker run -p 7860:7860 ghcr.io/charansripadi/tinyjudge.

Training

  • Base: answerdotai/ModernBERT-base, sequence classification (no / yes), one model for all three tasks.
  • Data v2: 2,806 training examples, generated from a simulated shop with rule-derived (exact) labels: correct and mistaken agent trajectories, template and LLM-written replies, one-fact corruptions.
  • 3 epochs, lr 3e-5, batch 16, fp16, max length 512. About 4 minutes on a free Colab T4.
  • Tracked with MLflow; the dataset manifest's SHA-256 is logged with every run.

Limitations

  • One domain. Trained on a demo online-shop support agent with six tools. Other domains need new data.
  • Small real test set. 38 to 85 examples per task; differences of a few points are within noise.
  • Label assumptions. Real grounded positives assume the agent's own reply matched its tool results; at least one such reply contained an invented timeline, so some "errors" are the model being right.
  • LLM baseline. Qwen 3.8 27B without reasoning. A stronger or reasoning model would likely do better on policy.
  • Cascade. Escalating policy hurt accuracy here, because the LLM is worse than TinyJudge on that task.
Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for CharanSripadi/tinyjudge

Finetuned
(1507)
this model