Configuration Parsing Warning:In adapter_config.json: "peft.base_model_name_or_path" must be a string

Gev (gemma4-e2b-gev)

Gev = Gemma + the Kev/Jev naming pattern. A System One decision model on google/gemma-4-E2B: one document (the state) and a set of typed questions in (Choice / Noul / Score), a calibrated probability distribution per question out, in one forward pass. No text is generated. It follows the architecture and recipe of Kev β€” an open reconstruction of TypeSafe's Jev β€” with Gemma 4 E2B in place of Qwen3.5 as the backbone, and serves the same /v1/systemone API.

Not affiliated with TypeSafe AI (Jev) or with the author of Kev. Gev is an independent model; it was not trained on any Jev output.

At a glance

On Kev's frozen suites, locked test read once, against the released Kev-0.8B scored with the same harness on the same machine (Kev's numbers reproduce its model card exactly):

Kev-0.8B (released) Gev paired Ξ”, 95% CI (record-clustered bootstrap)
out-of-domain accuracy (transfer-v4 test, 656) 0.697 0.726 +2.9 pp [βˆ’1.2, +7.0], 85 wins / 66 losses
in-distribution accuracy (decision-v7 test, 1,200) 0.838 0.839 +0.2 pp [βˆ’1.9, +2.3]
out-of-domain Brier / ECE 0.397 / 0.046 0.367 / 0.037
in-distribution Brier / ECE 0.231 / 0.019 0.229 / 0.023

Ahead on both suites, not significant at 95%. The checkpoint was chosen among 10 candidates on the development splits, so development numbers are optimistic; the test split above was read once, after selection.

Development splits (same items for every row):

out of domain (transfer-v4 dev) in distribution (decision-v7 dev) rule composition (in dist.) held-out rule structures
Kev-0.8B released 0.648 0.827 0.859 0.625
Kev-0.8B stage 1 (v7-base) 0.642 0.829 0.828 0.667
Gev 0.683 0.835 0.938 0.688

Largest per-source test differences vs Kev-0.8B: SST-5 +23.7 pp, TweetEval-offensive +8.7, held-out contrastive policies +8.7, IMDB +7.5; banking77 βˆ’8.7, MNLI βˆ’8.7, Yelp βˆ’3.7.

Model

  • Backbone: google/gemma-4-E2B @ d29ff6b45f081a49ee2733a859c9c9c2d95d1a6f, text decoder only β€” the vision and audio towers are dropped at load time. Frozen.
  • LoRA r=16, Ξ±=32 on q/k/v/o_proj and gate/up/down_proj (24.95M trainable parameters) + Kev's pointer head (query/key projections to 256 dims over the <decide> and option positions).
  • Delimiters: Gemma's reserved <unused0>…<unused4> (unforgeable from caller text); a <bos> is prepended, as Gemma expects.
  • Gemma 4's sliding-window layers cannot honour Kev's packed block-causal mask, so every question runs as its own causal row continuing from the (cached) state β€” the row form Kev uses for its hybrid Qwen3.5 bases. Training path vs cached serving path agree to max |Ξ”p| 1.1e-5 (fp32, 2,172-token state).
  • Temperature 1.74 stored in head.pt (argmax unchanged; KEV_TEMPERATURE=1.0 gives raw logits).

Training

Two stages on one L40S (~1.5 h + ~0.5 h):

  1. Kev's base recipe on the frozen suite evals/v7/decision-v7 (ten public classification sources + programmatic policy and rule-composition data): 2 epochs, lr 5e-5 (Kev's setting for its 4B/9B bases; 1e-4 over-fit here), effective batch 8 (4 Γ— 2 accumulation, gradient checkpointing), bf16 autocast with fp32 frozen weights, option permutation / none-of-the-above / distractor augmentation, none minimal pairs on 25% of Choice records. 70 of 12,576 records (all banking77, 77-way) exceed the training context under Gemma's tokenizer with Kev's admission headroom and are excluded.
  2. Delta fine-tune (Kev's --init_from … --data … --replay 2000 path): 7,200 freshly generated rule-composition records from kev.composition (the eight training shapes plus 120 new random rule structures; held-out and locked structures and the locked render style excluded; zero text overlap with any evaluation partition) mixed with 2,000 replayed suite records, lr 4e-5, 1 epoch.

Why stage 2: after stage 1 the model beat Kev on natural-language and knowledge sources (MMLU, emotion, Amazon, BoolQ) but trailed badly on programmatic rule composition, and only on accept cases (it defaulted to "deny"). More rule data fixed that. Hard-case mining (training only on the pairs the model got wrong) did worse than a random sample of the same size, and so was not used.

Temperature: fitted on decision-v7 development rows, the method Kev's shipped temperatures use (in distribution).

Usage

The adapter needs Kev plus the Gemma patch in this repo (kev-gemma4.patch: Gemma delimiters and <bos>, text-decoder extraction, row form for sliding-window backbones, a tokenizer-aware training admission filter, and the scripts used above).

git clone https://github.com/jaredpalmer/kev.git && cd kev
git checkout 0c142be
curl -L https://huggingface.co/sungwon1110/gemma4-e2b-gev/resolve/main/kev-gemma4.patch | git apply
uv sync --extra serve
KEV_DTYPE=bf16 uv run --extra serve python -m kev.serve --run sungwon1110/gemma4-e2b-gev --port 8009
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
  "state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
  "questions": {
    "department": {"type": "choice", "instructions": "Which team should handle this?",
                   "criteria": {"returns": "Exchanges, refunds, wrong or damaged items",
                                "shipping": "Delivery status, delays, lost packages",
                                "billing": "Charges, invoices, payment problems"}},
    "escalate":   {"type": "noul", "instructions": "Does this need urgent human attention?"}
  }}'

Speed (CPU)

8 threads on a shared Xeon Silver 4514Y (other VMs running; figures are relative), one question, same cores and alternating requests for every row:

input Kev-0.8B bf16 Gev bf16 Gev fp32
~60 tokens 0.68 s 0.58 s 1.49 s
~1.6k tokens 4.7 s 10.2 s 15.8 s
~6.5k tokens 26.1 s 68.4 s 97.8 s

On CPU use bf16 (AMX): fp32 is 1.4–1.6Γ— slower. Short inputs run at Kev-0.8B speed; long states are 2–2.6Γ— slower. On an L40S the full fp32 dev suites score in about two minutes.

Limitations

  • Trained on English data only. A small Korean probe was answered, but Korean was not evaluated systematically.
  • The win over Kev-0.8B is not statistically significant at 95% on either locked suite; one training seed.
  • Weaker than Kev-0.8B on banking77 (partly the 70 excluded training records) and MNLI.
  • Not trained for any operations-specific task; zero-shot incident-escalation judgments on an internal IT-operations probe were weak (near-chance ranking) and should not be relied on without task data.
  • Like Kev and Jev: a classifier over the options you give it β€” it cannot say anything outside them, and its probabilities are advisory outside the training distribution.

License

Apache-2.0, as are the base model (Gemma 4) and Kev.

Downloads last month
14
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for sungwon1110/gemma4-e2b-gev

Adapter
(40)
this model