klev-e4b

klev (KV + Jev) is a 4B decision model: given a typed question β€” a single choice, a noul (true/false) judgement, a score, or a multi-choice item β€” it returns calibrated probabilities over the options plus an explicit rejection channel. It is built on unsloth/gemma-4-e4b-it-unsloth-bnb-4bit, fine-tuned with QLoRA, and read out through a small pointer head instead of text generation. It runs 4-bit on a single 16 GB GPU.

It was built using Unsloth: the 4-bit QLoRA fine-tune, the delimiter embedding deltas, the pointer head and the teacher-KL cache all run through Unsloth's patched Gemma 4 stack (FastModel). See "Built with Unsloth" below, and the unsloth usage snippet under Usage.

The name comes from KV + Jev: a Kev-lineage pointer readout serving Jev's typed decision protocol. Working name during development was M5.

Built on Kev

klev-e4b is based on Kev and reuses part of its code. The record format, the pointer readout, the metrics and the decision-v7 data are Kev's (Apache-2.0, Jared Palmer), vendored rather than reimplemented so that klev stays byte-compatible with the Kev family β€” which is what makes the Kev-4B column in Results a like-for-like baseline:

from Kev reused here
kev/model.py (PointerHead) model/head.py β€” the q/k readout below, plus klev's own garbage candidate
kev/api.py, kev/data.py data/format.py β€” TypeSafe record format, render(), record builder
kev/metrics.py data/metrics.py β€” the scorer used for every number in Results (verbatim)
kev/suite.py + jaredpalmer/kev-suites data/suites.py β€” frozen, sha256-checked suites
kev.calibrate scripts/calibrate.py β€” the fitted pointer temperature in head.pt
decision-v7 the 15,576 training rows (the same partitions Kev-0.8B/4B/9B trained on)

Kev is Apache-2.0 and every vendored file carries a header naming its origin. No Kev model weights are used: the base is Gemma 4 E4B IT, the head and LoRA are trained fresh. What is klev's own is the Gemma base, the Unsloth QLoRA training stack, the KL anchor, the delimiter embedding deltas, the rejection channel and the stitch below.

Built with Unsloth

klev was built using Unsloth (2026.9.7), which is what makes the recipe fit one consumer GPU β€” 4-bit NF4 QLoRA of a 4B base, pointer head and cached-teacher KL on a single RTX 4080 16 GB.

stage Unsloth usage
base load unsloth.FastModel.from_pretrained(..., load_in_4bit=True)
fine-tune FastModel.get_peft_model for the rank-16 LoRA + use_gradient_checkpointing="unsloth", driven by a transformers.Trainer subclass (DistillTrainer)
trainable tokens delimiter rows via peft trainable_token_indices (Unsloth's Gemma 4 embedding path)
teacher KL cached base logits, read through Unsloth's patched hidden-state/logits access
inference / evals FastModel for every decision, chat, drift and stitch eval

Reproducing the run needs unsloth imported before transformers / peft / trl, because it patches them at import time:

pip install unsloth            # 2026.9.7
python -c "import unsloth, transformers, peft, trl"   # unsloth first

Files

file what it is
adapter/ rank-16 LoRA (alpha 32) + tokenizer with the 5 delimiter special tokens (<unused0>..<unused4>)
head.pt pointer head weights and the trainable delimiter-embedding deltas
train_config.json training hyperparameters (decision-v7, 2 epochs, KL anchor 0.3)
stitch/costbr-lora/ external third-party LoRA (CoSt-BR, pt-BR stance) trained on the same base
stitch/probe.pt 27-example nearest-class-mean probe in that LoRA's prompt space + T, beta, ext_weight
stitch/report.json stitch bench report (CoSt-BR, 718 rows)

This is not a chat model: decisions are read from the pointer head over the option/decision tokens. Reuse model/delimiters.py, model/head.py and eval/eval_decisions.py from the code repo to load and score it.

Architecture

  • Base: Gemma 4 E4B IT, 4-bit NF4 QLoRA, loaded and trained with Unsloth (FastModel).
  • Five Gemma reserved tokens (<unused0>..<unused4>) mark state / option / decision positions; their embedding rows are trained as deltas (no vocabulary resize).
  • Pointer head (vendored from Kev, see "Built on Kev"): q/k dot product over option boundary tokens, softmax over the K options plus a learned garbage candidate for rejection.
  • Trainable: 43.7M params (LoRA + deltas + head).

Training

  • decision-v7: 15,576 decision rows, 2 epochs, 8.5 h on one RTX 4080 16 GB (~6.5 s/step).
  • Objective: pointer CE over K+1 + KL anchor 0.3 to the frozen base on decision states.
  • The KL anchor keeps the representation base-aligned (drift: mean KL 1.864, 88.4 % next-token argmax agreement on held-out Alpaca prompts), which is what makes the model composable with external LoRAs (see below).

Results

Same records and scorer for all models. Best per row in bold.

suite klev-e4b Kev-4B Winnow-E4B
MMLU 0.657 0.725 0.701
ARC-Challenge 0.856 0.916 0.896
HellaSwag 0.752 0.800 0.795
Stanceosaurus ar / en / ru / es 0.421 / 0.431 / 0.437 / 0.596 0.428 / 0.485 / 0.580 / 0.655 0.449 / 0.443 / 0.583 / 0.653
CoSt-BR (majority 0.404) 0.400 0.323 0.415
in-distribution multilingual 0.785 0.794 0.823
OOD multilingual (XNLI/Belebele/XStory/PAWS-X) 0.751 0.756 0.850*

* Winnow answered 3,520/4,000 OOD rows. Per task: XNLI is flat (0.74–0.76), Belebele favours Winnow (0.877 vs 0.699), XStoryCloze favours klev (0.924 vs 0.874 Kev-4B / 0.922 Winnow).

Own decision suite: decision-v7 dev 0.860 (Brier 0.212), transfer-v4 dev 0.751 β€” identical on both loaders. Chat-mode retention stays level with the base (MMLU 0.322 β†’ 0.389).

Same checkpoint, two loaders

Every row above was re-scored with the plain loader (--backend plain: transformers + bitsandbytes, no unsloth) on the same prepared rows, through the same readout (eval/run_card_suites.py, docs/m15-plain-inference.md).

suite n unsloth plain delta choice agreement
decision-v7 dev 1,468 0.860 0.860 +0.000 100.0 %
transfer-v4 dev 764 0.751 0.751 +0.000 99.7 %
MMLU 989 0.657 0.659 +0.002 99.2 %
ARC-Challenge 1,000 0.856 0.856 +0.000 99.8 %
HellaSwag 1,000 0.752 0.753 +0.001 99.3 %
Stanceosaurus (4 langs) 5,923 0.471 0.471 +0.000 98.8 %
CoSt-BR 718 0.400 0.404 +0.004 98.5 %
OOD multilingual 4,000 0.751 0.750 -0.001 99.6 %
in-distribution multilingual 3,128 0.785 0.784 -0.001 99.6 %

Brier moves by ≀0.001 everywhere. Two options, same weights:

  • plain (the default install) β€” pip install "klev @ git+https://github.com/felipepenhorate/klev.git". torch + transformers + peft + bitsandbytes. This is what the second column above is.
  • unsloth β€” the stack the model was trained with: pip install "klev[train] @ git+...", then load_klev(ckpt, backend="unsloth"), KLEV_BACKEND=unsloth, or --backend unsloth. Needed for fine-tuning and teacher caching; for pure serving it only matters if you want the exact kernels the benchmark numbers were produced with.

The residual disagreement is kernels, not weights (unsloth's patched attention/quantised matmuls vs plain sdpa + bnb): at most 0.12 absolute probability on any single row, and no suite moves by more than 0.4 points.

Stitch: importing an external LoRA with ≀30 examples

Because of the KL anchor, a LoRA trained outside this fine-tune composes with klev without wrecking its decisions (decision-v7 dev 0.860 β†’ 0.855 with stitch/costbr-lora active). Its knowledge is linearly readable from the hidden state of its own prompt format, so fusing that readout into the pointer logits with two scalars imports it:

logits_final = logits_ptr(h_sys) + beta * log p_alp
p_alp = softmax(-d^2 / T)          # NCM in the external LoRA's prompt space
CoSt-BR (718 test rows) ext weight 0.5 ext weight 1.0
klev alone (pointer) 0.405 0.389
27-example probe, LoRA space 0.457 0.475
stitched klev 0.521 0.511
external LoRA alone (generation) 0.522 0.522

Latency (RTX 4080, 4-bit, single stream):

median p90
vanilla decision 205 ms 231 ms
stitched 456 ms 495 ms
overhead +251 ms (2.23Γ—) +264 ms

The overhead is one extra forward over the LoRA's prompt; the fusion itself is negligible. Probe hyperparameters: T = 344.5, beta = 4.0, ext_weight = 0.5 (a robust cell; the train-selected cell gives 0.521 β€” see stitch/report.json).

Install

# klev is not on PyPI: install from the repository
pip install "klev @ git+https://github.com/felipepenhorate/klev.git"

klev is the library: it carries the loader, the pointer readout, the record format and both usage paths, so a consumer does not clone this repo to use a checkpoint. For this base the version floor is load-bearing:

why
transformers >= 5.5.0 the gemma4 architecture does not exist before it. On 5.3.0 AutoConfig raises "Transformers does not recognize this architecture", which unsloth then misreports as a generic "not supported yet" ValueError β€” it reads like a missing model, not a version.
unsloth >= 2026.9.11 first release carrying gemma4 support.

This arm needs a 16 GB class card: Gemma 4 E4B in 4-bit plus the 262,144-row embed_tokens and embed_tokens_per_layer tables does not fit 8 GB, and the per-layer table is an nn.Embedding, which bitsandbytes never quantizes. If you want the same model on a small card, use lumierenoir/klev-0.8b instead.

Two console scripts come with it:

klev-system-one --ckpt lumierenoir/klev-e4b --request my_request.json   # no server
klev-serve       --ckpt lumierenoir/klev-e4b --port 8090                # optional HTTP

load_klev() reads the base model from the checkpoint's own adapter_config.json and selects the matching delimiters, so you do not have to pass base= or preset=.

Usage

klev is a library first. A decision model is small and stateless, so the normal way to use one is in-process: load it once, call it as many times as you like. No server, no port, no daemon. import unsloth must come before transformers / peft (it patches them at import time), which is why it is first.

import unsloth  # noqa: F401  β€” must precede transformers / peft
from klev import load_klev, answer      # pip install "klev @ git+https://github.com/felipepenhorate/klev.git"

model, tokenizer, head = load_klev("lumierenoir/klev-e4b")   # or a local dir with adapter/ + head.pt

result = answer(model, tokenizer, head, {
    "state": "Customer writes: my card keeps getting declined and the ATM refused it too.",
    "questions": {
        "is_card": {"type": "noul",   "instructions": "Is the card being declined?",
                    "criteria": {"false": "no, something else is the subject",
                                 "true": "yes, the card is being declined"}},
        "action":  {"type": "choice", "instructions": "What should the agent do first?",
                    "criteria": {"card_usage": "ask how the card was used",
                                 "atm_fault": "raise an ATM fault with the network",
                                 "fraud_check": "run a fraud check on recent transactions"}},
        "urgency": {"type": "score",  "instructions": "How urgent is this?",
                    "criteria": ["no action needed", "next working day", "same day", "immediate"]},
    }})
result["answers"]["action"]["choice"]   # the option key
result["answers"]["action"]["none"]     # rejection channel

none is the rejection channel β€” the mass on "none of these apply", renormalised off the K options so the returned distribution still sums to 1. The three question types are noul (criteria = {false, true}), choice (criteria = an option map) and score (criteria = an ordered level list). A request is validated by klev's own SystemOneRequest before any GPU work, and labels / target / _meta are ignored β€” so a labelled eval record can be passed straight through.

Temperature matters. answer(..., temperature=1.0) is the raw readout, which is what the Results table above scores. The served numbers use the temperature fitted by scripts/calibrate.py on this model's own calibration rows: 1.625 (Brier 0.212 / ECE 0.037 at coverage@5 % 0.722, versus raw 0.227 / 0.079 / 0.722). head.pt stores 1.0, because it is written during training and the head only tempers in eval mode, so a deployment has to supply the fitted value itself.

From a shell, still no server

klev-system-one --ckpt lumierenoir/klev-e4b                    # built-in demo
klev-system-one --ckpt lumierenoir/klev-e4b --request r.json    # your request
echo '{"state":"...","questions":{...}}' | klev-system-one --ckpt ... --request -

which ends with a plain-language readback β€” the quickest way to sanity-check a checkpoint.

If you would rather have HTTP

Worth it when several processes share one warm copy of the weights, or when the caller is not Python.

klev-serve --ckpt lumierenoir/klev-e4b --port 8090 --temperature 1.625
curl -s localhost:8090/v1/systemone -H 'content-type: application/json' -d @request.json
route
POST /v1/systemone {"state", "questions"} β†’ {"answers": {qid: typed answer}}
GET /v1/models loaded base, temperature, whether the rejection channel exists
GET /health {"ok": true, "model": ...}

Same routes and response shapes as kev's own server, so eval/eval_systemone_http.py scores it on the bench records unmodified. It is stdlib http.server, so there is no fastapi/uvicorn to install. One decision per request under a lock, since a forward is not re-entrant.

What was validated where. The code here is the same model/infer.py and serve/server.py that lumierenoir/klev-0.8b documents, and every check listed in docs/m14-usage.md was run against that arm β€” inline calls on all three question types, GET /health and GET /v1/models, POST /v1/systemone with the full TypeSafe shape, 400 on a malformed body, 404 on an unknown route, and 120 decision-v7 development rows through eval/eval_systemone_http.py with zero failures. This arm was not re-run: Gemma 4 E4B does not fit the 8 GB card the testing was done on. The code path is identical and load_klev reads the base from this checkpoint's own adapter_config.json, but treat the numbers above as the arm that was measured, not this one.

License and data

  • Base model: Apache-2.0 (google/gemma-4-E4B-it, unsloth/gemma-4-e4b-it-unsloth-bnb-4bit).
  • This release (adapter, head, probe): Apache-2.0.
  • Kev (github.com/jaredpalmer/kev, Apache-2.0, Jared Palmer) β€” klev is built on it and reuses its record format, PointerHead, metrics, suite loading, calibration procedure and decision-v7 data. See "Built on Kev" above.
  • Unsloth (github.com/unslothai/unsloth, Apache-2.0) β€” the training and inference stack this model was built with. See "Built with Unsloth" above.
  • Training data: decision-v7 records built from permissively licensed sets (Apache-2.0 / MIT / CC-BY-4.0), attribution manifest in the code repo.
  • stitch/costbr-lora was trained in training-conversational-stance/ on the CoSt-BR dataset (pt-BR Reddit conversations); check that dataset's terms before redistribution.

Code

Architecture, training, evals and the stitch experiment: github.com/felipepenhorate/klev (SPEC.md, docs/m5–m14). The library is pip install "klev @ git+https://github.com/felipepenhorate/klev.git"; klev-0.8b and its base/ arm document the same usage paths for a card that fits 8 GB.

Predictions do not need unsloth

Unsloth is a training dependency. The default install is torch + transformers + peft + bitsandbytes, and a published checkpoint is served through plain transformers β€” same accuracy (0.410 vs 0.410 on 195 CoSt-BR rows for klev-e4b, 0.317 vs 0.317 for klev-0.8b), 98 % identical choices. Unsloth is the train extra:

pip install "klev @ git+https://github.com/felipepenhorate/klev.git"            # predictions
pip install "klev[train] @ git+https://github.com/felipepenhorate/klev.git"     # + fine-tuning

Backend is explicit when you want the unsloth path: load_klev(ckpt, backend="unsloth"), KLEV_BACKEND=unsloth, or --backend unsloth on klev-system-one / klev-serve. See docs/m15-plain-inference.md.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support