Kev-9B, fine-tuned on browser decisions

AIMultiple fine-tuned Kev-9B (revision 2629c06) on browser-agent decisions for the decision models benchmark. On the benchmark's 50 browser tasks it completed 32, against 21 for stock Kev-9B run on the same day and hardware (exact McNemar, p = 0.003).

Stock Kev-9B This fine-tune
Benchmark: browser tasks completed 21/50 32/50
Held-out tasks from the training websites, completed 13/41 21/41
Held-out single decisions, same operation as the labeling model 62% (63/101) 94% (95/101)
Finished pages recognized as done 13/18 18/18

Like Kev-9B, this is a LoRA adapter plus a pointer head on Qwen/Qwen3.5-9B-Base (revision 68c46c4).

Usage

Tested with kev at commit 5e94a28.

uv run --extra serve python -m kev.serve --run AIMultiple/kev-9b-browser-ft --port 8008

The benchmark served it in fp32 with the reference kernels and an 8,192-token input limit. It was trained and scored on requests from the jev-ultrafast browser runtime: page text, a numbered list of visible controls, an operation question and target questions.

Results

Stock and fine-tuned Kev ran the benchmark's 50 browser tasks on the same H100 on the same day, one attempt per task, with the same runtime, limits and grader.

  • The fine-tune completed 12 tasks stock Kev failed and failed 1 that stock Kev completed. One of the 12, C02, was a browser timeout in the stock run rather than a model failure; without it the gain is 11 to 1 (p = 0.006).
  • Attempts ending blocked fell from 12 to 5, and attempts stopped at the action limit from 11 to 6. Declaring an unfinished task done rose from 4 to 7.
  • The held-out rows use tasks from the training websites that were kept out of training. The single-decision test replays 103 recorded decisions from 19 of those tasks; Kev could read 101 of them.

Training

  • Start point: jaredpalmer/kev-9b at revision 2629c06.
  • Labels: GLM-5.3-Flash (MIT license) ran browser tasks on eight websites that are not in the benchmark. The 184 successful training runs gave 606 decision records. A run counted as successful when the task check passed it; across all tasks, 11 more runs counted because they reached the right product page through site search and failed only the exact-URL check.
  • Recipe: kev.train at commit 5e94a28, with the settings of the producer's own delta fine-tune: LoRA rank 16, learning rate 2e-5, gradient accumulation 4, bf16, gradient checkpointing, p_none 0.1, p_none_distract 0.12, p_distract 0.15. The full settings are in training_config.json.
  • Our changes: 2 epochs instead of 1, batch size 1 instead of 2, no replay of the producer's training records (their run replayed 2,000), and training input limits raised to the serving limits. kev.train otherwise encodes at 384 state, 1,024 branch and 2,048 packed tokens.
  • Training took about 16 minutes on one H100 (training_metrics.json).

Limitations

  • Each task was attempted once, and the 50 tasks come from nine websites.
  • Labels come from one model on eight websites, and the training data is not published.
  • head.pt carries a temperature of 1.0, where Kev-9B ships 2.30, so the probabilities are not calibrated. The benchmark runtime acts only on the top-ranked option, and a temperature does not change the ranking, so the results above are unaffected.

License and changes

Apache-2.0 for the adapter and head, the same as Kev-9B. The base model, Qwen3.5-9B-Base, is also Apache-2.0. Changes from Kev-9B: new adapter weights and pointer head from the fine-tune described above.

Checksums (sha256)

  • adapter_model.safetensors: 3386d5af6c24c599325a55db25a21f7c556682b8cc41b46a6545e8a874e35b99
  • head.pt: daf245f4f77de2d50089ab421d3f6c8fdaa32195f26bb6e46f180b457350331c
Downloads last month
19
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AIMultiple/kev-9b-browser-ft

Finetuned
(2)
this model