Instructions to use AIMultiple/kev-9b-browser-ft with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use AIMultiple/kev-9b-browser-ft with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Kev-9B, fine-tuned on browser decisions
AIMultiple fine-tuned Kev-9B (revision 2629c06) on browser-agent decisions for the decision models benchmark. On the benchmark's 50 browser tasks it completed 32, against 21 for stock Kev-9B run on the same day and hardware (exact McNemar, p = 0.003).
| Stock Kev-9B | This fine-tune | |
|---|---|---|
| Benchmark: browser tasks completed | 21/50 | 32/50 |
| Held-out tasks from the training websites, completed | 13/41 | 21/41 |
| Held-out single decisions, same operation as the labeling model | 62% (63/101) | 94% (95/101) |
| Finished pages recognized as done | 13/18 | 18/18 |
Like Kev-9B, this is a LoRA adapter plus a pointer head on Qwen/Qwen3.5-9B-Base (revision 68c46c4).
Usage
Tested with kev at commit 5e94a28.
uv run --extra serve python -m kev.serve --run AIMultiple/kev-9b-browser-ft --port 8008
The benchmark served it in fp32 with the reference kernels and an 8,192-token input limit. It was trained and scored on requests from the jev-ultrafast browser runtime: page text, a numbered list of visible controls, an operation question and target questions.
Results
Stock and fine-tuned Kev ran the benchmark's 50 browser tasks on the same H100 on the same day, one attempt per task, with the same runtime, limits and grader.
- The fine-tune completed 12 tasks stock Kev failed and failed 1 that stock Kev completed. One of the 12, C02, was a browser timeout in the stock run rather than a model failure; without it the gain is 11 to 1 (p = 0.006).
- Attempts ending blocked fell from 12 to 5, and attempts stopped at the action limit from 11 to 6. Declaring an unfinished task done rose from 4 to 7.
- The held-out rows use tasks from the training websites that were kept out of training. The single-decision test replays 103 recorded decisions from 19 of those tasks; Kev could read 101 of them.
Training
- Start point:
jaredpalmer/kev-9bat revision2629c06. - Labels: GLM-5.3-Flash (MIT license) ran browser tasks on eight websites that are not in the benchmark. The 184 successful training runs gave 606 decision records. A run counted as successful when the task check passed it; across all tasks, 11 more runs counted because they reached the right product page through site search and failed only the exact-URL check.
- Recipe:
kev.trainat commit5e94a28, with the settings of the producer's own delta fine-tune: LoRA rank 16, learning rate 2e-5, gradient accumulation 4, bf16, gradient checkpointing, p_none 0.1, p_none_distract 0.12, p_distract 0.15. The full settings are intraining_config.json. - Our changes: 2 epochs instead of 1, batch size 1 instead of 2, no replay of the producer's training records (their run replayed 2,000), and training input limits raised to the serving limits.
kev.trainotherwise encodes at 384 state, 1,024 branch and 2,048 packed tokens. - Training took about 16 minutes on one H100 (
training_metrics.json).
Limitations
- Each task was attempted once, and the 50 tasks come from nine websites.
- Labels come from one model on eight websites, and the training data is not published.
head.ptcarries a temperature of 1.0, where Kev-9B ships 2.30, so the probabilities are not calibrated. The benchmark runtime acts only on the top-ranked option, and a temperature does not change the ranking, so the results above are unaffected.
License and changes
Apache-2.0 for the adapter and head, the same as Kev-9B. The base model, Qwen3.5-9B-Base, is also Apache-2.0. Changes from Kev-9B: new adapter weights and pointer head from the fine-tune described above.
Checksums (sha256)
adapter_model.safetensors:3386d5af6c24c599325a55db25a21f7c556682b8cc41b46a6545e8a874e35b99head.pt:daf245f4f77de2d50089ab421d3f6c8fdaa32195f26bb6e46f180b457350331c
- Downloads last month
- 19