jeff-adapter-code-gguf
The LoRA GGUF of jeff-adapter-code (Jeff v1.3), for llama.cpp.
Jeff-Code: information-gathering steps ahead of the coding model. In the Jeff-Code agent, takes the information-gathering steps ahead of Qwen3.8-27B (reads, listings, searches) and hands over when unsure.
This repository holds only the adapter. Load it on top of the base GGUF from mstrasser/jeff-base-gguf; there is no separate full model per adapter.
It works on either base format: Q8_0 (recommended, within 0.4 points of full precision on every adapter test) or Q4_K_M (within 0.4 points).
Files
| File | What it is |
|---|---|
loras/code.gguf |
The LoRA (169 MB), the same file for every base format. It stays at full precision; quantisation applies to the base weights only |
models/code.jeff.json |
What a client needs: the answer codes and their token ids, the prompt layout (live-last) and the temperature fitted for each format |
code.jeff.json has a temperature for each format (temperature_by_format: f16, q8_0, q4_k_m); use the one for
your base format, so the probabilities stay calibrated.
Results
The same rows, through llama.cpp: the base GGUF plus this LoRA GGUF, with the temperature refitted for each format. Running Jeff with llama.cpp
| Test set | Full precision | Q8_0 | Q4_K_M |
|---|---|---|---|
development |
90.4% 路 0.023 | 90.4% 路 0.030 | 91.0% 路 0.029 |
Run-time rule: the same action (act on the top option at 0.40, otherwise hand over). The GGUF takes the same decision as full precision on 99.9% of the 991 development rows at Q8_0 and 97.9% at Q4_K_M.
Full precision: full precision from the trainer's own evaluation on the development split.
More on the adapter's card: mstrasser/jeff-adapter-code.
How to run it
Use llama.cpp commit cb7934c52ca8710994b2ecc19775ebefcfdb8d01 or newer: it needs the qwen35 architecture and LoRA
on the output layer.
hf download mstrasser/jeff-base-gguf --local-dir jeff-gguf
hf download mstrasser/jeff-adapter-code-gguf --local-dir jeff-gguf
cd jeff-gguf
llama-server -m models/base-q8_0.gguf -c 8192 -np 1 --lora-init-without-apply --lora loras/code.gguf
Jeff is not a chat model. For each decision you run one forward pass over the prompt and read the probabilities of the answer codes; you never sample text.
- Build the prompt exactly as Jeff does (the question, the state, the options as answer codes, and the changing
state field under "Latest", then the chat template with thinking off). The Jeff repository builds it for you
(
jeff.model.decision_messages). - Tokenize without a beginning-of-sequence token, with special tokens parsed.
- Run one forward pass over the whole prompt from an empty state, with this adapter active. Clear the state between prompts: Qwen3.5 has recurrent layers, so a reused state would carry over.
- Read the probabilities. Take the logits of the first N answer-code tokens (
token_ids[:N], N = the number of options), divide by the temperature for this adapter and your format, and apply a softmax over those N.
With llama-server, each decision is one POST /completion with "n_predict": 1, "cache_prompt": false, a lora
list that names every loaded adapter (scale 1 for this one, 0 for the rest; an adapter left out of the list keeps
scale 1.0), a logit bias of +1000 on every answer-code token of the request, and n_probs set to the number of
options. From the llama.cpp library: load the base once with llama_model_load_from_file, this adapter once with
llama_adapter_lora_init, and per request call llama_set_adapters_lora with scale 1, clear the memory, decode once
and read llama_get_logits_ith(ctx, -1).
The full guide, with an example prompt and request: Running Jeff with llama.cpp.
Data and licence
Adapter licence: Apache-2.0.
Qwen3.5-0.8B notice: these weights were modified from Qwen3.5-0.8B by the Jeff project: jeff-base is a fine-tune of Qwen3.5-0.8B, and this adapter was trained on top of it. Qwen3.5-0.8B is Copyright 2026 Alibaba Cloud and licensed under the Apache License, Version 2.0; a copy of that licence is in LICENSE.
To confirm: that the licences of the session data and of Terminal-Bench 2.0 allow training and publishing the adapter.
It was trained on:
ukisai/Qwen3.8-27B-multi-turn-agent-sft (about 14,400 public Qwen3.8-27B agent sessions). Licence: To confirm (licence not confirmed yet) 路 Made by Qwen3.8-27B (the sessions); labels built by code
Jeff's option lists are rebuilt from each transcript; the label is the step Qwen took next.
To confirm: the licence of the ukisai data set on Hugging Face
openguardrails Terminal-Bench sessions of Qwen3.8-27B. Licence: To confirm (licence not confirmed yet) 路 Made by Qwen3.8-27B (the sessions); labels built by code
Public sessions, replayed in each task's container.
To confirm: the licence of the openguardrails Terminal-Bench sessions
Terminal-Bench 2.0 tasks (outside the 40 evaluation tasks and their four near-twins). Licence: To confirm (licence not confirmed yet) 路 Not made by a model
The tasks our own Jeff-Code sessions and the replays run on.
To confirm: the Terminal-Bench 2.0 licence
Our own Jeff-Code sessions with Qwen3.8-27B. Licence: Built for this adapter (made for this adapter) 路 Made by Qwen3.8-27B (local), the sessions; labels built by code
Qwen3.8-27B in Jeff-Code (bash only); Jeff logs its full option list before every turn without acting.
Links
- The adapter at full precision: mstrasser/jeff-adapter-code (revision v1.3)
- The base GGUF: mstrasser/jeff-base-gguf
- Adapter card: mstrasser/jeff-adapter-code
- Running Jeff with llama.cpp: jeffhub.ai/docs/llama-cpp
Jeff is an independent project. It uses the same request format as Jev but is not affiliated with or endorsed by TypeSafe, the makers of Jev.
- Downloads last month
- -
We're not able to determine the quantization variants.