Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

jeff-adapter-code-router

Jeff-Code router: should the coding model think hard on this turn?. In the Jeff-Code agent, picks how hard Qwen3.8-27B should think on each turn (off, low, medium, xhigh), so easy turns run fast.

A LoRA adapter for jeff-base v1.3, a small open decision model (a fine-tune of Qwen3.5-0.8B). You send a situation (the state) and questions with named options; Jeff returns a calibrated probability for every option from one forward pass, with no generated text to parse. One Jeff server loads the base once and any number of adapters beside it; each request picks an adapter by name ("model": "code-router").

llama.cpp: serve this adapter on the Q8_0 base only. Its choices are close calls, so at Q4_K_M it changes about 6% of its thinking on/off decisions (Q8_0: 0.8%).

Results

  • 62.4% vs 62.8%: pass rate of Jeff-Code vs Qwen3.8-27B alone; paired difference −0.2 points (95% interval −2.6 to +2.1), 1,242 paired tasks on 6 benchmarks
  • 47% faster (32% less time) per task: 0.68× Qwen alone's time on average (95% interval 0.64–0.72; median task 0.70×)
  • 14% less total time over all tasks combined (0.86×, 0.80–0.93)
  • 92% / 68%: offline accuracy on held-out scoring tasks: code (steps) / code-router (thinking)

Jeff-Code runs two adapters on the fixed Jeff v1.3 base: code takes the information-gathering steps (step threshold 0.40), and code-router decides whether Qwen thinks hard on a turn (thinking off unless P(xhigh), out of the four levels off/low/medium/xhigh, ≥ 0.6). The baseline is Qwen3.8-27B alone in the same Jeff-Code build with every Jeff feature off, thinking at full on every turn and no thinking limit, which is how plain Pi runs it. Each task ran in both settings side by side, at the same time on the same Qwen server, and every comparison is paired by task. Only tasks Jeff never saw in training were used: the held-out splits of six benchmarks.

Benchmark Paired tasks Qwen alone Jeff-Code Difference, points (95% interval) Time per task
SWE-bench Verified 486 70.6% 70.8% +0.4 (−3.5 to +4.3) 0.63× (0.57–0.69)
SWE-rebench (2 rounds) 370 58.9% 58.1% −0.8 (−5.7 to +3.5) 0.66× (0.61–0.73)
Terminal-Bench Pro (2 rounds) 195 61.2% 62.8% +1.5 (−4.6 to +7.7) 0.64× (0.55–0.76)
Terminal-Bench 2.0 (40 tasks, 3 attempts each) 108 75.9% 70.0% −4.6 (−12.1 to +3.7) 0.96× (0.78–1.16)
SkillsBench 42 28.6% 31.0% +2.4 (−11.9 to +16.7) 0.91× (0.68–1.20)
Harbor Index 41 12.2% 9.8% −2.4 (−12.2 to +7.3) 0.71× (0.51–0.99)
All six, pooled 1,242 62.8% 62.4% −0.2 (−2.6 to +2.1) 0.68× (0.64–0.72)
  • No benchmark shows a clear pass-rate difference: every interval includes zero. SWE-rebench and Terminal-Bench Pro ran twice; their two rounds are combined, with intervals computed over both. Time per task is the geometric mean of the per-task time ratios (Jeff-Code's time divided by Qwen alone's); below 1 is faster. The pass rates count every finished session; the difference counts only tasks finished in both settings, so it is not exactly the gap between the two pass rates.
  • Total time drops less than time per task: in about 5% of tasks, Jeff-Code runs more than 30 minutes longer than Qwen alone, because it keeps going where Qwen alone gives up after a few minutes. The good news is, sometimes that pays off: in those tasks Jeff-Code solved 26 to Qwen's 24.
  • Thinking off on every turn (with the same safeguards and thinking limit) is faster still but clearly worse: −7.6 points (−10.6 to −4.5), −13.5 on Terminal-Bench 2.0. A less cautious router threshold (0.7) loses 4.5 points. Jeff's decisions are what keep the quality.
  • Terminal-Bench (original) and Terminal-Bench Science also ran, but both settings solve 0% of their tasks, so they are left out. 27 task pairs that hit an infrastructure failure twice are left out for both sides; about 10 long re-runs were still running when these numbers were taken.

The agent: the Jeff-Code repository (github.com/firelex/jeff-code). Offline accuracy reported by the Jeff-Code session: practically the same as the full fine-tune compared offline (step 92%, router 68% for both). Numbers from the evaluation report of 2026-10-05 18:23. Source: results/sources/v1.3/jeff-code-owner-supplied.json in the JeffHub repository (supplied by the maintainers).

Jeff-Code: the whole story

What it is. Jeff-Code is a coding agent based on Pi, with two Jeff v1.3 adapters trained specifically for Qwen3.8-27B: this one and its partner (code for steps, code-router for thinking). We forked Pi because its extension framework doesn't currently let a fast decision model sit deep enough inside the agent loop. Besides being useful to people who run Qwen3.8-27B locally as their daily coding model, it is an experiment: how far can a small, fast "System 1" model go inside a coding agent?

How it works. Jeff makes two kinds of decisions around every Qwen turn, each by a small adapter on the same Jeff base, in about 0.2 s each:

  1. Jeff works ahead of Qwen (code). If it can take the next information-gathering step itself (read a file, list a folder, search the code, check which tools are installed), it does. It picks the tool first and then its argument, and can take several steps in a row while it is confident. Qwen then starts its turn with those results already in front of it, rather than spending a slow turn fetching them. It can also run the tests or a build, repeat Qwen's last command and, with the run-approval setting the evaluation used, run a script Qwen wrote or install a missing package. Writing and editing files always stay with Qwen, and whenever Jeff is unsure, it hands over to Qwen. So Jeff does not only pick a tool; it also fills in the tool's argument.
  2. Jeff decides whether Qwen needs to think hard on its next turn (code-router). Thinking stays off unless Jeff is confident the turn needs it. In the evaluation, about three quarters of Qwen's turns ran with thinking off.

Because Jeff returns a calibrated probability for every choice, each behaviour is controlled by a single setting: the step threshold (0.40) and the thinking threshold (0.6).

How we trained it: Jeff predicts what Qwen would do next.

  • Steps: each training label is simply the step Qwen actually took next. If Jeff can take that step early, Qwen gets the result without spending a turn on it. These labels are built from Qwen sessions by code, with no other model involved.
  • Thinking: each recorded Qwen turn at full thinking was asked again with thinking off, then low, then medium. The label is the cheapest level whose action was as good as the original, or full thinking if none was. "As good" is decided by code wherever possible (the same kind of step on the same target); otherwise Qwen3.8-Max, with thinking off, judges whether the cheaper step would serve the task just as well at that moment. About 27% of turns didn't need thinking at all.
  • These labels are deliberately strict, which made the router cautious, and that is what preserved the pass rate. We also tried the looser question, "Is the full-thinking step materially better?". On a single turn the judge can't reliably tell thinking-off from a second full-thinking answer either, so a router trained that way would switch thinking off almost everywhere, and the thinking-off run shows what that costs over a whole task.

Adapters, not a fine-tune. Everything measured above used adapters: small LoRA files on the fixed Jeff v1.3 base. We also trained one full fine-tune to make both decisions and compared it offline; its accuracy was practically identical (step 92%, router 68% for both). So we release the adapters, which keep the multi-adapter design without giving up accuracy.

What changed from Pi.

  • Thinking per turn: Pi only switches thinking on or off, so Qwen always thought at its highest level (in our tests, Qwen's low and medium levels thought about as long as the highest one, so they saved no time). Jeff-Code sets Qwen's thinking level for each turn, fixed or decided by Jeff.
  • Safeguards for thinking off: a loop guard catches repeated or near-identical actions up to six steps back (a file write only counts as progress if it changes the file); a caught repeat is thrown away and that turn is asked again with full thinking, and near-identical outputs or two failed commands in a row send the next turn to full thinking. The thinking-off comparison had exactly the same safeguards, the same thinking limit and the same escalations (about 3% of its turns ended up thinking); the only difference from Jeff-Code is Jeff's decisions.
  • Runaway cut-off: if Qwen's thinking or text keeps repeating itself, the reply is stopped and asked again with full thinking. This was on in every run, including the baseline (it caught 2 replies there).
  • Thinking limit: at 8,000 thinking tokens, Qwen answers from what it has thought so far. That also rescues replies that would otherwise hit the 32K output limit, which ends a Pi session. The baseline ran without it, as plain Pi does.
  • Jeff steps: before each Qwen turn, Jeff-Code builds a menu of concrete next steps from what is already known, and Jeff takes them when it is confident.
  • A pool of Jeff servers, one per GPU, keeps each decision at about 0.2 s, and every Qwen request and Jeff decision is logged.

With llama.cpp (GGUF)

The same rows, through llama.cpp: the base GGUF (mstrasser/jeff-base-gguf) plus this adapter's LoRA GGUF (mstrasser/jeff-adapter-code-router-gguf), with the temperature refitted for each format. Running Jeff with llama.cpp

Serve this adapter on the Q8_0 base only. Its choices are close calls, so at Q4_K_M it changes about 6% of its thinking on/off decisions (Q8_0: 0.8%).

Test set Full precision Q8_0 Q4_K_M
development 64.4% · 0.063 64.4% · 0.063 63.8% · 0.067

Run-time rule: the same thinking on/off decision (on when P(xhigh) ≥ 0.6). The GGUF takes the same decision as full precision on 99.2% of the 987 development rows at Q8_0 and 94.3% at Q4_K_M.

Full precision: full precision from the trainer's own evaluation on the development split.

When to use it

  • You run the Jeff-Code agent with Qwen3.8-27B locally and want hard thinking only on the turns that need it.
  • You want a single setting (the thinking threshold) that trades speed for care.

When not to use it

  • You use another coding agent or another large model. The router was measured only with Qwen3.8-27B in Jeff-Code.

How to use it

The adapter runs with Jeff's server on the jeff-base v1.3 base. Adapter serving arrives with the next Jeff release; until then, these commands need the feat/lora branch of firelex/jeff.

git clone https://github.com/firelex/jeff && cd jeff
uv sync --no-default-groups --extra lora          # add --extra cuda on NVIDIA GPUs, --extra mac on Apple silicon
uv run --no-default-groups hf download mstrasser/jeff-base --revision v1.3 --local-dir checkpoints/jeff-base
uv run --no-default-groups hf download mstrasser/jeff-adapter-code-router --revision v1.3 --local-dir adapters/code-router
JEFF_CHECKPOINT=checkpoints/jeff-base JEFF_ADAPTERS=adapters/ PORT=8765 \
  uv run --no-default-groups jeff-serve          # on a Mac, add JEFF_BACKEND=mlx

Every folder in adapters/ is served under its folder name; add or replace adapters while the server runs with curl -X POST http://localhost:8765/v1/adapters/reload. Each adapter records the exact base it was trained on, and the server refuses an adapter trained on a different one, so this adapter loads only on jeff-base v1.3 (a v1.2 adapter does not load on v1.3). For llama.cpp, use mstrasser/jeff-adapter-code-router-gguf.

Request format

State (the situation), in this order:

Key Changes per request What it holds
task no The coding task, as given to the agent.
recent_steps yes The last few steps (each tool call and its output), trimmed to fit.

Questions:

  • thinking (choice): How hard Qwen should think on its next turn. Options: Four levels - off, low, medium and xhigh

Rules:

  • Jeff-Code builds these requests itself; you do not write them by hand.
  • In the measured runs, thinking stays off unless the probability of xhigh, out of all four options, is at least the thinking threshold (0.6).
  • The request format may still change before release.

General rules for every request: the request format guide.

Example

The request below is also in this repository as example.json.

{
  "model": "code-router",
  "state": {
    "task": "The unit tests in tests/test_parser.py fail after the last change. Find out why and fix it.",
    "recent_steps": "bash: sed -n 30,60p src/parser.py -> (the date-parsing function, with a regular expression that no longer matches two-digit years)"
  },
  "questions": {
    "thinking": {
      "type": "choice",
      "instructions": "How hard should the coding model think before its next turn?",
      "criteria": {
        "off": "Off: the next step is routine; answer without thinking",
        "low": "Low: think briefly",
        "medium": "Medium: think the step through",
        "xhigh": "Extra high: the next step needs careful reasoning"
      }
    }
  }
}
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d @adapters/code-router/example.json

The answer holds a probability for each option of each question. A recorded response from the v1.3 adapter is not published yet.

Files

  • adapter_model.safetensors, adapter_config.json: the LoRA weights (PEFT format);
  • readout.safetensors: the adapter's own readout over the answer codes;
  • decision_config.json: answer codes, temperature, prompt layout and the checksum of the base it was trained on;
  • example.json: the example request above.

adapter_config.json and decision_config.json name the base as mstrasser/jeff-base, revision v1.3; the server checks the base by the checksum of its weights.

Training

Base mstrasser/jeff-base, revision v1.3 (a fine-tune of Qwen3.5-0.8B)
Prompt layout live-last: the fixed part of the request first, the changing state field last
LoRA GGUF for llama.cpp mstrasser/jeff-adapter-code-router-gguf
  • 1.3.0 (2026-10-05): First release, trained on Jeff v1.3 with the live-last prompt layout. Used in Jeff-Code's measured runs with the thinking threshold at 0.6.

Data card

Self-reported. The numbers come from the adapter’s own maintainers and have not been re-run by anyone else. What the levels mean

  • Test set: not attached yet
  • QA report: not available here yet; it will be added once sanitised

How the test set was held out. Whole tasks are held out: Jeff-Code was measured only on tasks never used for training. SWE-bench Verified ran in full (all 500 tasks; none of its repositories were used for training). For the benchmarks also used in training, the tasks were split and every held-out task was run: Terminal-Bench 2.0 (40 frozen tasks; four near-twins of them were also kept out of training), SWE-rebench (189), Terminal-Bench Pro (100), SkillsBench (44) and Harbor Index (41). A leak check compares every training task with the evaluation tasks.

Training data. Training data not published.

How it works: before each of Qwen's turns, the router gives a probability for each of four thinking levels - off, low, medium and xhigh. In the measured runs, thinking stays off unless the probability of xhigh, out of all four, is at least the thinking threshold (0.6). The code adapter takes the routine information-gathering steps; together they are Jeff-Code's two decisions.

Why it matters: thinking off throughout is faster still but clearly worse (7.6 points lower pass rate, 95% interval −10.6 to −4.5; −13.5 on Terminal-Bench 2.0); the router's decisions are what keep the quality.

How the labels are made: each recorded Qwen3.8-27B turn (recorded at xhigh) was asked again with exactly the same request at thinking off (temperature 0), then low, then medium (normal sampling), stopping at the first level whose action was good. The label is that level, or xhigh if none was.

What counts as good: first a code rule - the same intent as the recorded xhigh action, meaning the same kind of step and the same target per command. Otherwise Qwen3.8-Max (hosted, Alibaba Cloud DashScope; thinking off, temperature 0) answered yes or no to: 'Would Step 2 serve the task as well as Step 1 at this moment?', seeing the task and the last 3 steps. Inline scripts, program runs, file writes and edits, and final answers always went to that judge.

Sessions recorded at medium thinking were left out.

Label shares in training: off about 27%, low about 8%, medium about 5%, xhigh about 60%.

In the evaluated runs, low and medium were never used: thinking was off unless the probability of xhigh was at least 0.6.

The row counts are still to be added to this card.

The terms of the hosted model providers are being checked for training and publication use.

Questions longer than 8,192 tokens are cut to fit, keeping the head and tail of the state; the question and its options are never cut. The same rule applies at run time.

Jeff-Code's measured results come entirely from LoRA adapters on the fixed Jeff v1.3 base - this one and code - loaded beside the other adapters on one jeff-base.

Data and licence

Adapter licence: Apache-2.0.

Qwen3.5-0.8B notice: these weights were modified from Qwen3.5-0.8B by the Jeff project: jeff-base is a fine-tune of Qwen3.5-0.8B, and this adapter was trained on top of it. Qwen3.5-0.8B is Copyright 2026 Alibaba Cloud and licensed under the Apache License, Version 2.0; a copy of that licence is in LICENSE.

To confirm: that the licences of the session data and of the benchmark tasks allow training and publishing the adapter.

It was trained on:

  • Recorded Qwen3.8-27B sessions (at xhigh), re-asked at lower thinking levels. Licence: To confirm (licence not confirmed yet) · Made by Qwen3.8-27B (the re-asked answers); Qwen3.8-Max (hosted, Alibaba Cloud DashScope), for the yes-or-no judgements the code rule could not make

    Sessions recorded at medium thinking were left out. The re-asked answers come from Qwen3.8-27B; where the code rule could not decide, Qwen3.8-Max judged whether the lower-level step was as good.

    To confirm: the licence of the source sessions

Limitations

  • Tied to jeff-base v1.3. It will not load on any other base or version; the server checks the base weights' checksum.
  • Jeff chooses between the options you give it. It does not write text or reason in several steps.
  • Calibration was fitted on this adapter's own calibration rows. On very different data, check it again.
  • Everything listed under When not to use it above.

Links

Jeff is an independent project. It uses the same request format as Jev but is not affiliated with or endorsed by TypeSafe, the makers of Jev.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mstrasser/jeff-adapter-code-router

Adapter
(30)
this model