CloudSync Pro support agent — Qwen2.5-1.5B

Qwen2.5-1.5B-Instruct taught to work a first-line customer-support queue for a fictional product: search a knowledge base, answer from what it returns, and hand the conversation to a human when it belongs to one.

Two checkpoints of one lineage, each in its own folder:

folder stage dev score
01M2HK1C6W92J2VEWDFNX6GEXK/train-60/ SFT, 60 steps from the base model 87.19% ± 4.63
01M2HP4M1G955GXWV3XBFR5YT6/train-110/ + GRPO, 50 steps — use this one 96.07% ± 2.69

The untrained base model scores 11.79% ± 4.47 on the same exam.

The task

The agent answers one customer message with two tools:

  • search_kb(query) — retrieves from an 8-entry knowledge base (plans and prices, refund window, file size limits, version history, offline mode, SSO, sync troubleshooting, storage).
  • escalate_to_human(reason) — hands over, stating one of six categories: money, legal, dataloss, cancel, angry, kbgap.

A reply is scored by a grader, not a judge model. Escalation rows score on routing and on naming the right category; answer rows score on stating the facts the question asked for (recall) and not reciting facts it didn't ask for (precision). Both halves divide by how much was said, so hedging — escalating everything, or reciting the whole knowledge base — scores near the floor.

Results

Measured on 200 held-out tasks written to the same blueprint as the training data but with disjoint text, greedy decoding, scored by the environment's own grader.

dev score routing correct category correct answer rows
Qwen2.5-1.5B-Instruct (base) 11.79% ± 4.47 — — —
+ SFT (step 60) 87.19% ± 4.63 194 / 200 95 / 96 0.794
+ GRPO (step 110) 96.07% ± 2.69 196 / 200 96 / 97 0.961

Intervals are 95% (1.96·√(p(1−p)/n), n=200). Paired against the base model task by task, the GRPO checkpoint is +0.86 ± 0.048, fixing 172 tasks and breaking none.

The base model almost never used the tools at all — 1 search and 0 escalations across 200 tasks — so SFT accounts for the tool use, and GRPO for the accuracy of what gets said: answer rows rise from 0.794 to 0.961.

How it was trained

SFT — 1,931 demonstrations, published as monte-inc/cloudsync-support-sft: each one replayed through the environment's real grader and kept only at a perfect score, none of them drawn from either exam. 60 steps, batch 32, lr 5e-6, sequences to 1,024 tokens.

GRPO — 50 steps on 1,921 prompts (the same customer messages, labels only, no written answers). 32 prompts per step, 16 attempts each, temperature 1.0, sequences to 2,048 tokens, reward from the same grader that scores the exams. Mean reward rose from 0.75 to 0.94.

Trained with NeMo-RL v0.7.0, rollouts through NeMo Gym, on one H100.

Serving

The checkpoints are standard Hugging Face folders. Tool calls come back in Qwen's Hermes format, so a server needs the matching parser:

vllm serve monte-inc/qwen2.5-1.5b-cloudsync-support \
  --revision main \
  --enable-auto-tool-choice --tool-call-parser hermes

To serve one checkpoint directly, download its folder first:

from huggingface_hub import snapshot_download

path = snapshot_download(
    "monte-inc/qwen2.5-1.5b-cloudsync-support",
    allow_patterns="01M2HP4M1G955GXWV3XBFR5YT6/train-110/**",
)

Requires transformers ≥ 5. config.json was written by transformers 5, which stores the RoPE base inside rope_parameters. transformers 4.x doesn't read that key: it silently falls back to rope_theta=10000, where Qwen2.5 needs 1000000. The model still answers, but degrades badly — in our own testing it put a stray brace in nearly every tool call, and 2 calls in 40 parsed. On an older stack, add "rope_theta": 1000000.0 at the top level of config.json.

Provenance

Every number above comes from a recorded run in the monte ledger, experiment support6: Baselines 01M2HJW87K6BSZK2WFGSKPYEN8 and 01M2HJYZNC8BHB9NX34ZHPE7NK, SFT 01M2HK1C6W92J2VEWDFNX6GEXK, GRPO 01M2HP4M1G955GXWV3XBFR5YT6, evals 01M2HKGG8Y108GVS0B4S8RAD7T and 01M2HY6KKVEQZ4GYNZPPE1Z9N3. Exams, curriculum and training pile are frozen by content hash; the grader that scored the exams is byte-identical to the one that paid the training reward.

Optimizer state for each checkpoint, for resuming training, is in a companion repo: …-resume-<run>-train<step>. Everything from this work sits together in the 2026-09 CloudSync Pro Support collection.

Limitations

  • CloudSync Pro is not a real product. The knowledge base, the customers and their problems are synthetic, written for this environment. The agent knows 8 facts and nothing else.
  • The dev exam shares a blueprint with the training data. Its rows were written without sight of it and screened for near-duplicates, but it flatters the model compared with genuinely new traffic.
  • The held-out sealed exam has not been read on either checkpoint. The base model scores 4.63% ± 2.91 on it; there is no trained score to report yet.
  • Six categories, one message. It is not trained on multi-turn conversations, and it is a 1.5B model.
  • Intended for research and demonstration, not for handling real customers.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for monte-inc/qwen2.5-1.5b-cloudsync-support

Finetuned
(1884)
this model

Dataset used to train monte-inc/qwen2.5-1.5b-cloudsync-support

Collection including monte-inc/qwen2.5-1.5b-cloudsync-support

Evaluation results

  • Grader reward (routing x facts/category) on support_escalation_dev@3 (held-out dev exam, 200 tasks)
    self-reported
    0.961