Instructions to use per2021/kev-2b-qwen3.5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use per2021/kev-2b-qwen3.5 with PEFT:
from peft import PeftModel from transformers import AutoModel base_model = AutoModel.from_pretrained("Qwen/Qwen3.5-2B-Base") model = PeftModel.from_pretrained(base_model, "per2021/kev-2b-qwen3.5") - Notebooks
- Google Colab
- Kaggle
Kev-2B (Qwen3.5 base), unofficial
A 2B-parameter Kev decision model: a rank-16 LoRA adapter and a pointer head on
Qwen/Qwen3.5-2B-Base, trained with the Kev-0.8B release recipe. It answers yes/no (noul), multiple-choice (choice)
and rating (score) questions about a text with calibrated probabilities, through Kev's /v1/systemone API.
This is not an official Kev release. Jared Palmer did not train, review or endorse it. It is an independent derivative built with the unmodified Kev code, to fill the gap between Kev-0.8B and Kev-4B on laptops (Apple Silicon, MLX).
Results
Test partitions of the public Kev suites, read once after the checkpoint was selected on development data only. Every
model was served the same way: kev.serve on an Apple M4 (24 GB), MLX backend, bf16, each checkpoint at the temperature
stored in its head.pt, scored with kev.benchmark --remote ... --allow-test. On this harness the official checkpoints
reproduce their published numbers: Kev-4B transfer-v4 test 0.843 (published 0.838) and documents-v1 test 0.903 (0.904);
Kev-0.8B transfer-v4 test 0.695 (0.697).
| Accuracy (test) | Kev-0.8B | Kev-2B (this) | Kev-4B |
|---|---|---|---|
| transfer-v4 (six never-trained sources, 656 q) | 0.695 | 0.796 | 0.843 |
| hard-v1 (long policies, trade-offs, probability, multi-hop, dates, judging, abstention; 1,088 q) | 0.669 | 0.699 | 0.801 |
| devtools-v1 (code review, commits, flaky tests, prompt safety; 1,073 q) | 0.637 | 0.691 | 0.755 |
| documents-v1 (real CFPB complaint narratives up to ~7k tokens; 936 q) | 0.851 | 0.859 | 0.903 |
Paired, record-clustered bootstrap (kev.metrics.paired_bootstrap, 1,000 resamples, macro), accuracy delta with 95% CI.
devtools-v1's two duplicated ids are dropped on both sides, as in the author's comparisons.
| Kev-2B minus | transfer-v4 | hard-v1 | devtools-v1 | documents-v1 |
|---|---|---|---|---|
| Kev-0.8B | +0.109 [+0.064, +0.160] | +0.030 [+0.000, +0.060] | +0.036 [+0.006, +0.059] | +0.010 [β0.023, +0.042] |
| Kev-4B | β0.056 [β0.102, β0.007] | β0.102 [β0.134, β0.071] | β0.074 [β0.107, β0.049] | β0.068 [β0.104, β0.034] |
| same recipe without the documents (skills-only delta) | +0.021 [β0.001, +0.045] | β0.019 [β0.037, β0.002] | β0.014 [β0.035, +0.003] | +0.152 [+0.100, +0.197] |
Calibration (served, test): ECE 0.046 on transfer-v4 and on documents-v1. The out-of-fold development check of the fitted temperature (T = 2.14) is separated from the raw one: ECE 0.080 β 0.017 [0.013, 0.038].
Size and speed on an Apple M4 (MLX, bf16): 5.1 GB resident, median 140β160 ms for a short text with 1β3 questions (Kev-0.8B 62 ms / 4.3 GB, Kev-4B ~420 ms / 9.1 GB). On an Apple M1 (16 GB, macOS 15.8), served from the Hub: same transfer-v4 development accuracy (0.752), 5.1 GB resident (7.0 GB peak while loading), median 278 ms, 764 benchmark requests in 4.5 minutes.
Summary. Kev-2B is significantly better than Kev-0.8B on out-of-domain decisions and developer tooling, marginally better on hard reasoning (the interval touches zero), ties it on long real documents, and sits 6β10 points below Kev-4B at 55% of its memory and about a third of its latency.
Candidates and selection
Two joint deltas were trained with the same recipe. They differ only in the documents labelled: 4,846 records when the
teacher budget ran out, and 5,182 after it was completed. The final checkpoint was chosen on development data only,
before any test read of it. Development accuracy was tied (mean over the four suites 0.734 vs 0.733;
decision-v7 0.852 vs 0.853), and the complete-data run was kept. Both candidates' test partitions were read once.
For transparency, the other candidate scored 0.788 / 0.720 / 0.701 / 0.861 on the four suites above. Against it, the
final checkpoint is +0.014 [β0.004, +0.033] on transfer-v4, β0.021 [β0.040, β0.002] on hard-v1, β0.013 [β0.031,
β0.001] on devtools-v1 and +0.000 [β0.012, +0.011] on documents-v1. That is run-to-run variation of a point or two, in
both directions; neither dominates.
Limitations
- Still overconfident on some unanswerable questions (e.g. 0.87 that a customer "was not angry" when the text says nothing). Use a confidence threshold and send the rest to a person.
- Weakest areas: sarcasm and fine-grained ratings in short texts, and hard multi-step reasoning (0.70 vs 0.80 for the 4B).
- Trained and evaluated in English. Spanish inputs worked in a small informal check but are not benchmarked.
How it was trained
The Kev-0.8B release lineage (night2-du base + round-15 joint delta), applied to the 2B base. Every stage ran
kev.train from the Kev repository through modal_app.py::study on one H100, bf16 autocast with fp32 master weights,
LoRA r=16 Ξ±=32 on attention, MLP and DeltaNet projections, and option permutation / none-of-the-above augmentation
(p_none_pair 0.25).
| Stage | Init | Data | Epochs | lr | Batch Γ accum | State context | Wall |
|---|---|---|---|---|---|---|---|
| 1. Base | Qwen/Qwen3.5-2B-Base@b1485b2 (pointer head from scratch) |
decision-v7 train (12,576 records) |
2 | 5e-5 | 8 Γ 1 | 384 | 35 min |
| 2. Dates / unknowable delta | stage 1 | night2/dates_unknowable (1,425) + 2,000 replay |
1 | 3e-5 | 8 Γ 1 | 384 | 6 min |
| 3. Joint documents + skills delta | stage 2 | documents (5,182, see below) + hard-v1 train (6,000) + devtools-v1 train (5,320) + 6,000 replay |
1 | 2e-5 | 2 Γ 4 | 7,552 | 66 min |
Stage 1 ran two learning rates (1e-4 and 5e-5, seed 0); 5e-5 was kept. It is the better arm on both the in-distribution
development partition (0.852 vs 0.849), which is the author's selection rule, and transfer-v4 development (0.729 vs 0.703).
The temperature was then fitted on the decision-v7 development rows with scripts/calibrate_checkpoint.py and written
into head.pt.
Data
decision-v7,night2/dates_unknowable,devtools-v1train: the Kev suites as published (hashes match the manifests).hard-v1train: regenerated withscripts/build_hard_v1.py; byte-identical to the manifest sha256.- Documents:
documents-v1train is not published, so it was rebuilt with the repository's own tooling. Candidates come fromscripts/build_documents_v1.py(same 5,994 CFPB complaint narratives, same stratification). Labels come fromscripts/label_documents_v1.pywith the author's two open-weight teachers, DeepSeek V3.2 and Qwen3-235B-A22B thinking (-2507), called through OpenRouter instead of the Vercel AI Gateway, with an unchanged prompt and parser. The filter isfreeze_documents_v1.train_split(a question is kept only when both teachers chose the consumer's own label). The result is 5,182 records / 7,474 questions, against the author's 5,219 / 7,488 (42 questions lacked a teacher answer, against 23 for the author). No training text appears indocuments-v1development or test (checked by normalised text hash). Teacher cost: $17.51. - No Jev output was used anywhere. Training labels come only from open-weight teachers or programmatic solvers, as in Kev's data policy.
Deviations from the official process
- The base size (2B) has no official recipe. Learning rates interpolate between the 0.8B and 4B recipes and were not tuned.
- Stage 1 used two learning-rate arms instead of several seeds.
- Documents were relabelled (same teachers, same prompt and filter), so the kept set differs in detail from the author's (0.2% fewer questions).
- Hardware: H100 instead of H200 for the joint delta. Evaluation ran in bf16 on MLX (Apple M4) instead of fp32. On this harness the official checkpoints reproduce their published test numbers (table above).
- The author's private held-out panels (
documents-v2,breadth-v1) could not be evaluated. There is no measurement on tool use, narrative reasoning, chess and the otherbreadth-v1areas.
Use
git clone https://github.com/jaredpalmer/kev.git && cd kev && uv sync --extra serve
uv run --extra serve python -m kev.serve --run per2021/kev-2b-qwen3.5 --port 8009
On Apple Silicon the server uses the MLX backend automatically. The calibrated temperature is stored in head.pt.
License
Apache-2.0 for the adapter and head; the Qwen3.5 base is Apache-2.0; training datasets carry their own licenses (see the
Kev-0.8B model card for the decision-v7 sources; CFPB narratives are US government work).
- Downloads last month
- 40
Model tree for per2021/kev-2b-qwen3.5
Base model
Qwen/Qwen3.5-2B-Base