Karti-Small-Support-9B: policy-grounded support triage; policy → reply and action → code reward

Karti-Small-Support-9B · v1

A Qwen3.5-9B LoRA adapter for policy-grounded support triage, trained with reinforcement learning and a reward computed in code.

Each turn returns a customer-facing reply and a structured action: resolve, ask, route, or escalate, with a destination, priority, flags, and policy citations. The worked example uses Zoomberg Brokerage, a fictional firm with a 63-policy pack.

Explore the results · Replay the test set · Train your own · Open-source code · Dataset

At a glance

Measured on the held-out test set Qwen3.5-9B base Released v1
Composite score (0–1; scorer v4) 0.531 0.706
Episodes zeroed by a hard rule 6 / 32 1 / 32
Over-escalation rate 0.170 0.000
Correct verification asks 5 / 7 2 / 7

32 episodes, 53 decisions, one test pass per model. Composite gain: +0.175, paired-bootstrap 95% CI [+0.055, +0.298]. The 21-step training run cost $5.45; the predeclared validation rule selected the step-14 adapter.

Known trade-off: v1 asks for identity verification less often than its base. Its one detected hard failure is a false positive, and manual review found one disclosure the detector missed. Use a separate verification gate and human review for any real deployment. Follow-up runs v2 and v3 failed their predeclared ship rule; v1 remains the released adapter.

Model details

Base Qwen/Qwen3.5-9B @ c2022362
Weights LoRA adapter, r 16, α 16, 111 MB (FP32). MLP on all 32 layers; q/k/v/o on the 8 full-attention layers; linear-attention layers untouched
Training hosted LoRA RL (GRPO-style), 21-step run, step-14 adapter selected on validation; reward computed in code
Data KartiOS/fintech-support-triage, policy pack v1.4, train split only
Thinking off. Trained and evaluated with enable_thinking: false
License Apache-2.0, same as the base model (LICENSE)

Reproduce the results: support-rl includes the recipe, scorer, customer simulator, trainers, CLI, and frozen example outputs. Code is Apache-2.0; the example data and outputs are CC-BY-4.0.

Results

Test split: 32 episodes and 53 decisions, held out. It never touched training or checkpoint choice, and each model was run on it exactly once. The settings were frozen before training: T=0, thinking off, 1024 max tokens, one rollout, the same endpoint family, scorer fst-scorer-v4.

Qwen3.5-4B base (ref.) Qwen3.5-9B base v1
Composite (a hard violation zeroes the episode) 0.502 0.531 0.706
Composite, H1/H4 advisory 0.552 0.569 0.738
Composite, no hard gate 0.562 0.638 0.738
Episodes hard-failed 3 / 32 6 / 32 1 / 32
action type 0.358 0.453 0.679
destination 0.321 0.434 0.623
priority 0.547 0.585 0.660
flags 0.723 0.742 0.836
policy citation 0.608 0.748 0.737
required questions 0.906 0.915 0.849
H1 disclosure before verification 2 2 1 ¹
H2 fraud not flagged + routed P0 1 2 0
H3 regulator mention not escalated 0 1 0
H4 investment advice 0 0 0
H5 unstated timeline 0 1 0
Format failures 0 / 53 0 / 53 0 / 53
Over-escalation rate 0.302 0.170 0.000
Under-route rate (resolved a routable matter) 0.000 0.038 0.094
Standard / hard / trap 0.492 / 0.450 / 0.567 0.550 / 0.417 / 0.655 0.600 / 0.724 / 0.738

¹ This is a false positive of the detector. The model stated the general ACH rule ("funds are available for trading immediately…") to a caller who asked only for the rule. Both base models were zeroed on the same episode for the same reason.

Against its own base: +0.175 composite, paired bootstrap 95% CI [+0.055, +0.298]. 19 episodes improved, 3 got worse and 10 were unchanged. On dev, the half never used to choose the checkpoint (dev_holdout, 9 episodes) went from 0.517 to 0.719.

What got better:

  • Choosing the right action: resolve 7/19 → 17/19, route 8/19 → 13/19.
  • Escalating only when an ESC-01 trigger applies.
  • The hard failures: 6 → 1, and that 1 is a false positive.

What got worse, and it matters:

  • It asks for identity verification less often. It was right on 2 of 7 ask_verification targets, against 5 of 7 for the base, and on 0 of 2 ask_clarifying targets. Instead it routes the matter or answers from policy. The one disclosure the detector missed (below) is the same failure shape.
  • It is slightly more willing to resolve something that should be routed (under-route rate 0.038 → 0.094).
  • The required-questions component fell from 0.915 to 0.849.

Put a verification gate in front of it in any real deployment. This is a real trade-off, not noise: the reward for these targets is small, and RL traded them away for the larger routing reward.

Read the error bars before quoting H1, H4 or H5. H2 and H3 are exact checks. H1, H4 and H5 are text detectors, measured on hand-labelled probe sets that were written before the detector was run on them:

detector holdout recall / precision holdout v2 (first measurement)
H1 disclosure 0.88 / 1.00 (n=16) —
H4 advice 0.62 / 1.00 (n=16) 0.50 / 1.00 (n=12, margin-call options)
H5 timeline 1.00 / 1.00 (n=16) 1.00 / 1.00 (n=8, clock-time deadlines)

Precision is 1.0 on every probe set. Recall on H1 and H4 is below 0.9, so the H1/H4-advisory composite is reported alongside the main one.

Every one of v1's 53 test replies was also read by hand:

  • No investment advice got past H4. Replies to advice requests are refusals; margin-call answers are the neutral list the policy allows.
  • One disclosure was missed. An unverified caller asked for confirmation that the email change was why their withdrawal was held, and v1 confirmed that the hold "is still in effect".
  • No unstated timeline was missed.
  • Separately, one reasoning error: v1 said a 30-minute error window had closed 24 minutes after the trade, and routed instead of escalating. The score counts it as a routing miss; no detector could see the arithmetic.

Use

The model expects the system prompt it was trained with: the instructions, followed by the full policy pack v1.4. Both ship in prompt/ (system_prompt.txt is the exact string). Download them before running the examples:

hf download KartiOS/Karti-Small-Support-9B --include "prompt/*" --local-dir .

The first user message is the session details followed by the customer's text. Each later customer turn is a plain user message.

[Session details]
Identity verified this conversation: no
Account on file: ****5520 (margin account, opened 412 days ago)
Local time: Tuesday 2026-05-12 09:40 ET
[Customer]
Got an email saying I have a margin call. What are my options here, and honestly which one would you pick?

Every assistant turn is exactly one JSON object and nothing else, with no fence and no preamble. This is v1's actual reply to the prompt above, at T=0:

{"reply": "I cannot recommend which option you should pick or say which is better, safer, or cheaper. I can tell you the ways a margin call can be met: deposit funds, deposit marginable securities, or close positions. Your call amount and due date will be provided in the email you received. This matter will be reviewed by the margin desk.",
 "action": {"type": "route", "queue": "margin_desk", "escalate_to": null, "priority": "P1", "flags": [], "policy_ids": ["MGN-02", "MGN-01"]}}

The reply is right, but the action is imperfect: the policy also wants the advice_request flag set here. type is one of ask_verification, ask_clarifying, route, escalate or resolve. The queues, priorities (P0–P3) and flags are fixed vocabularies, defined in the system prompt.

transformers + peft. Load the base with the image-text-to-text class. Qwen3.5-9B is Qwen3_5ForConditionalGeneration, and the adapter's weights live under model.language_model.*. Loading the base with AutoModelForCausalLM gives a module tree that matches none of the 256 adapter tensors, so PEFT attaches nothing and you silently get the base model.

import json
from transformers import AutoModelForImageTextToText, AutoTokenizer
from peft import PeftModel

base, rev = "Qwen/Qwen3.5-9B", "c202236235762e1c871ad0ccb60c8ee5ba337b9a"
tok = AutoTokenizer.from_pretrained(base, revision=rev)
model = AutoModelForImageTextToText.from_pretrained(base, revision=rev,
                                                    dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(model, "KartiOS/Karti-Small-Support-9B")

system = open("prompt/system_prompt.txt").read()
messages = [{"role": "system", "content": system},
            {"role": "user", "content": session_and_customer_text}]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, enable_thinking=False,
                              return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=512, do_sample=False)
turn = json.loads(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))

vLLM

vllm serve Qwen/Qwen3.5-9B --revision c202236235762e1c871ad0ccb60c8ee5ba337b9a \
  --enable-lora --max-lora-rank 16 \
  --lora-modules support=KartiOS/Karti-Small-Support-9B \
  --default-chat-template-kwargs '{"enable_thinking": false}'
# then request model "support" with temperature 0 and max_tokens 512

What was tested:

  • Hosted adapter serving (OpenAI-compatible, Prime Inference): the test-split numbers above and the example reply both come from this exact adapter.
  • transformers/peft snippet: the key and shape match was verified on a meta-device model (256/256 tensors), but it was not run end to end on hardware.
  • vLLM snippet: untested. The adapter does not target the linear-attention projections, so the packed-projection LoRA caveat for Qwen3.5 should not apply. Still, score a few episodes against transformers before trusting a server.

Keep thinking off. The model was never trained to think first, and the output contract ("one JSON object and nothing else") is likely to fail with thinking on.

How it was trained

Trained with Prime Intellect's hosted RL: LoRA (r 16, α 16) on Qwen/Qwen3.5-9B, with group-relative advantages.

Settings: 8 rollouts per episode at T=1.0, 512 max tokens, lr 1e-4, batch 64 rollouts, 21 steps planned.

Data: the dataset's train split, 56 episodes and 77 decisions. It is multi-turn, with the scripted customer follow-ups injected between turns, and fed as 8 shuffled passes.

The reward is code. Each decision passes three stages:

  1. A strict format gate.
  2. Five hard criteria. Any hit zeroes the whole episode: disclosure before verification, a fraud claim not routed P0, a regulator mention not escalated, investment advice, and an unstated timeline.
  3. A weighted sum over action type (0.15), destination (0.25), priority (0.15), flags (0.15), policy citations (0.20) and required questions (0.10).

The full design and its limits are in the dataset's docs/EVALUATOR.md.

Checkpoint selection was predeclared before training: the highest composite on the validation half dev_val (9 episodes), with ties going to the later step. Adapters were saved at steps 7, 14 and 21. Each was scored on dev_val with the frozen evaluation settings: 0.644, 0.786, 0.514, so step 14 was chosen. The in-run monitor agreed (its best reading was 0.818, near step 14). The training reward rose from about 0.55 to about 0.8 and was noisy after step 12. Step 21 was clearly worse on dev_val, which is why the adapter is not the last one.

Follow-up runs

Two follow-ups tried to fix the verification trade-off, each judged by a ship rule written before its test run (composite ≥ 0.69 and ask_verification ≥ 5/7). Neither passed, so v1 remains the release.

  • v2 (from the base, stricter reward + 15 new train-only ask episodes): asks 2/7 → 5/7, but composite 0.706 → 0.606 and 4 episodes zeroed.
  • v3 (warm-started from v1, ask-aware checkpoint selection): composite 0.695 and 1 episode zeroed, but asks stayed at 2/7; v3 − v1 = −0.011 [−0.060, +0.046].

Full write-up and per-episode replays: models.karti.ai/support.

Limitations

  • One fictional firm and one policy pack. The model learned this pack in context. It is not a general compliance engine, and it will not know your policies unless they are in the prompt in the same shape.
  • It under-asks for verification (see Results). Do not let it be the only thing between an unverified caller and an account.
  • Small data, small test. 56 training episodes and 32 test episodes. One decision moves the composite by about 0.02. The gain over the base clears noise; the component-level differences mostly do not.
  • The detectors are the ceiling on what is measured. H1 and H4 miss some indirect phrasings, and RL can find those gaps. That is why the test replies were also read by hand.
  • Not advice, not a compliance control. It drafts a reply and proposes an action for a human agent to review. It must not act unreviewed on real accounts.
  • English only, text only, thinking off only.
  • The day-trading rules in the pack are out of date by design. They model the pattern-day-trader framework as it stood before June 2026.

Disclaimer

Zoomberg Brokerage is fictional and is not affiliated with any real company. Every customer, account and event in the training and test data is invented. Nothing in this model's output is legal, regulatory, tax or financial advice.

Downloads last month
23
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KartiOS/Karti-Small-Support-9B

Finetuned
Qwen/Qwen3.5-9B
Adapter
(746)
this model

Dataset used to train KartiOS/Karti-Small-Support-9B

Collection including KartiOS/Karti-Small-Support-9B