Instructions to use sungwon1110/gemma4-e2b-gev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use sungwon1110/gemma4-e2b-gev with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.base_model_name_or_path" must be a string
Gev (gemma4-e2b-gev)
Gev = Gemma + the Kev/Jev naming pattern. A System One decision model on google/gemma-4-E2B: one document (the state) and a set of typed questions in
(Choice / Noul / Score), a calibrated probability distribution per question out, in one forward pass. No text is
generated. It follows the architecture and recipe of Kev β an open
reconstruction of TypeSafe's Jev β with Gemma 4 E2B in place of Qwen3.5 as the backbone, and serves the same
/v1/systemone API.
Not affiliated with TypeSafe AI (Jev) or with the author of Kev. Gev is an independent model; it was not trained on any Jev output.
At a glance
On Kev's frozen suites, locked test read once, against the released Kev-0.8B scored with the same harness on the same machine (Kev's numbers reproduce its model card exactly):
| Kev-0.8B (released) | Gev | paired Ξ, 95% CI (record-clustered bootstrap) | |
|---|---|---|---|
| out-of-domain accuracy (transfer-v4 test, 656) | 0.697 | 0.726 | +2.9 pp [β1.2, +7.0], 85 wins / 66 losses |
| in-distribution accuracy (decision-v7 test, 1,200) | 0.838 | 0.839 | +0.2 pp [β1.9, +2.3] |
| out-of-domain Brier / ECE | 0.397 / 0.046 | 0.367 / 0.037 | |
| in-distribution Brier / ECE | 0.231 / 0.019 | 0.229 / 0.023 |
Ahead on both suites, not significant at 95%. The checkpoint was chosen among 10 candidates on the development splits, so development numbers are optimistic; the test split above was read once, after selection.
Development splits (same items for every row):
| out of domain (transfer-v4 dev) | in distribution (decision-v7 dev) | rule composition (in dist.) | held-out rule structures | |
|---|---|---|---|---|
| Kev-0.8B released | 0.648 | 0.827 | 0.859 | 0.625 |
Kev-0.8B stage 1 (v7-base) |
0.642 | 0.829 | 0.828 | 0.667 |
| Gev | 0.683 | 0.835 | 0.938 | 0.688 |
Largest per-source test differences vs Kev-0.8B: SST-5 +23.7 pp, TweetEval-offensive +8.7, held-out contrastive policies +8.7, IMDB +7.5; banking77 β8.7, MNLI β8.7, Yelp β3.7.
Model
- Backbone:
google/gemma-4-E2B@d29ff6b45f081a49ee2733a859c9c9c2d95d1a6f, text decoder only β the vision and audio towers are dropped at load time. Frozen. - LoRA r=16, Ξ±=32 on
q/k/v/o_projandgate/up/down_proj(24.95M trainable parameters) + Kev's pointer head (query/key projections to 256 dims over the<decide>and option positions). - Delimiters: Gemma's reserved
<unused0>β¦<unused4>(unforgeable from caller text); a<bos>is prepended, as Gemma expects. - Gemma 4's sliding-window layers cannot honour Kev's packed block-causal mask, so every question runs as its own causal row continuing from the (cached) state β the row form Kev uses for its hybrid Qwen3.5 bases. Training path vs cached serving path agree to max |Ξp| 1.1e-5 (fp32, 2,172-token state).
- Temperature 1.74 stored in
head.pt(argmax unchanged;KEV_TEMPERATURE=1.0gives raw logits).
Training
Two stages on one L40S (~1.5 h + ~0.5 h):
- Kev's base recipe on the frozen suite
evals/v7/decision-v7(ten public classification sources + programmatic policy and rule-composition data): 2 epochs, lr 5e-5 (Kev's setting for its 4B/9B bases; 1e-4 over-fit here), effective batch 8 (4 Γ 2 accumulation, gradient checkpointing), bf16 autocast with fp32 frozen weights, option permutation / none-of-the-above / distractor augmentation, none minimal pairs on 25% of Choice records. 70 of 12,576 records (all banking77, 77-way) exceed the training context under Gemma's tokenizer with Kev's admission headroom and are excluded. - Delta fine-tune (Kev's
--init_from β¦ --data β¦ --replay 2000path): 7,200 freshly generated rule-composition records fromkev.composition(the eight training shapes plus 120 new random rule structures; held-out and locked structures and the locked render style excluded; zero text overlap with any evaluation partition) mixed with 2,000 replayed suite records, lr 4e-5, 1 epoch.
Why stage 2: after stage 1 the model beat Kev on natural-language and knowledge sources (MMLU, emotion, Amazon, BoolQ) but trailed badly on programmatic rule composition, and only on accept cases (it defaulted to "deny"). More rule data fixed that. Hard-case mining (training only on the pairs the model got wrong) did worse than a random sample of the same size, and so was not used.
Temperature: fitted on decision-v7 development rows, the method Kev's shipped temperatures use (in distribution).
Usage
The adapter needs Kev plus the Gemma patch in this repo (kev-gemma4.patch: Gemma delimiters and <bos>,
text-decoder extraction, row form for sliding-window backbones, a tokenizer-aware training admission filter, and the
scripts used above).
git clone https://github.com/jaredpalmer/kev.git && cd kev
git checkout 0c142be
curl -L https://huggingface.co/sungwon1110/gemma4-e2b-gev/resolve/main/kev-gemma4.patch | git apply
uv sync --extra serve
KEV_DTYPE=bf16 uv run --extra serve python -m kev.serve --run sungwon1110/gemma4-e2b-gev --port 8009
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"returns": "Exchanges, refunds, wrong or damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems"}},
"escalate": {"type": "noul", "instructions": "Does this need urgent human attention?"}
}}'
Speed (CPU)
8 threads on a shared Xeon Silver 4514Y (other VMs running; figures are relative), one question, same cores and alternating requests for every row:
| input | Kev-0.8B bf16 | Gev bf16 | Gev fp32 |
|---|---|---|---|
| ~60 tokens | 0.68 s | 0.58 s | 1.49 s |
| ~1.6k tokens | 4.7 s | 10.2 s | 15.8 s |
| ~6.5k tokens | 26.1 s | 68.4 s | 97.8 s |
On CPU use bf16 (AMX): fp32 is 1.4β1.6Γ slower. Short inputs run at Kev-0.8B speed; long states are 2β2.6Γ slower. On an L40S the full fp32 dev suites score in about two minutes.
Limitations
- Trained on English data only. A small Korean probe was answered, but Korean was not evaluated systematically.
- The win over Kev-0.8B is not statistically significant at 95% on either locked suite; one training seed.
- Weaker than Kev-0.8B on banking77 (partly the 70 excluded training records) and MNLI.
- Not trained for any operations-specific task; zero-shot incident-escalation judgments on an internal IT-operations probe were weak (near-chance ranking) and should not be relied on without task data.
- Like Kev and Jev: a classifier over the options you give it β it cannot say anything outside them, and its probabilities are advisory outside the training distribution.
License
Apache-2.0, as are the base model (Gemma 4) and Kev.
- Downloads last month
- 14
Model tree for sungwon1110/gemma4-e2b-gev
Base model
google/gemma-4-E2B