Gev E4B

Google's Gemma 4 E4B, retrained to answer questions about a text with probabilities: which team should handle a ticket, how urgent it is, whether a rule allows a case. It is the default model of Gev, which serves it through an HTTP API compatible with Jev's /v1/systemone.

The model is meant to be run by Gev, which reads the probability of each answer at one position instead of generating text. It is not tested as a chat model.

Files

File Use
model.safetensors, config.json, tokenizer files the weights, bf16, for building a Mev engine on a Mac (scripts/build-engine.sh gev-e4b) or for your own tools
gev-e4b-q8_0.gguf for llama.cpp, 8.0 GB; Gev's Docker image downloads it
gev-e4b-q4_0.gguf for llama.cpp on smaller GPUs, 5.2 GB, laid out like Google's QAT Q4_0 GGUF

Retraining moved the weights off the grid Google's quantization-aware training had put them on, so Q4_0 costs this model more than it costs the original. On the 316 example requests of Gev's playground, the most probable answer through llama.cpp matched the one from the unquantized weights for 99.1 % of the questions with Q8_0 and 93.9 % with Q4_0 (Google's E4B at Q4_0: 96.2 %).

A ready-built engine for Apple's M4 is in hilman2/gev-e4b-mev.

Results

Measured on a Mac mini with an M4 Pro and 48 GB, through Gev's Mev engine.

Gev E4B Gemma 4 E4B Gemma 4 26B-A4B
Correct decisions in 7,671 test cases 88.8 % 83.6 % 88.2 %
Requests per second, 8 at a time 4.54 2.52

License

Apache License 2.0, as Gemma 4.

Downloads last month
85
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hilman2/gev-e4b

Quantized
(57)
this model
Finetunes
1 model