Cactus Hybrid — Gemma 4 E2B (GGUF)

A small, on-device model is fast and private, but sometimes wrong. At Cactus we post-train models to know when they are wrong: we ship probes inside the checkpoint that score every answer with a confidence between 0 and 1, returned as structured data (never parsed out of the answer text). Answer on-device when confidence is high; re-route to a bigger model when it's low:

if confidence < 0.85:
    answer = ask_a_bigger_model(prompt)

This repo holds GGUF builds of Cactus-Compute/gemma-4-e2b-it-hybrid for llama.cpp.

Benchmarks

Gemma 4 E2B Hybrid, the smallest Gemma model, matches Gemini 3.1 Flash-Lite on most benchmarks by routing only 15–35% of queries to Flash-Lite and running the rest itself:

Benchmark Handoff to match Flash-Lite (FP16) At 4-bit At 3-bit
ChartQA 15–20% 25–30% 40–50%
MMBench 30–35% 40–45% 50–55%
LibriSpeech 25–30% 35–40% 55–65%
GigaSpeech 30–35% 40–45% 50–55%
MMAU 30–35% 35–40% 50–55%
MMLU-Pro 45–55% ~90% n/a

Quantisation quality is measured on Cactus Quants, which performs well at uniform quantization; developers are encouraged to benchmark Unsloth, GGUF, and MLX quantization independently.

Quickstart

The gemma-4-e2b-it-hybrid architecture is not yet in upstream llama.cpp. Run these files with a build that includes the Cactus patch series — on unpatched llama.cpp they fail to load with "unknown model architecture" by design. Build the patched server once:

git clone https://github.com/cactus-compute/cactus-hybrid && cd cactus-hybrid
./patches/llama.cpp/install.sh && rehash   # clones the pinned tag, applies the patches, builds

Then serve and query it like any llama-server:

llama-server -hf Cactus-Compute/gemma-4-e2b-it-hybrid-GGUF:Q4_K_M --jinja
curl -s http://localhost:8080/v1/chat/completions \
  -d '{"messages":[{"role":"user","content":"What is the capital of France?"}],"max_tokens":512}' \
  | jq '{answer: .choices[0].message.content, confidence}'

Chat-completions responses (and the final SSE chunk when streaming) carry a top-level "confidence" field.

Files

file quant size notes
gemma-4-e2b-it-hybrid-f16.gguf F16 9.31 GB closest to the bf16 reference
gemma-4-e2b-it-hybrid-Q4_K_M.gguf Q4_K_M 3.43 GB recommended for consumer hardware

The probe head (11 probe.* tensors) is stored in F32 in all quants — only the trunk is quantized.

Calibration note

Quantized trunks shift the layer-28 activations the probe reads, moving confidences downward relative to the bf16 reference (measured mean drift: F16 −0.07, Q4_K_M −0.10; easy-vs-hard ordering fully preserved). If you use aggressive thresholds, calibrate per quant; the 0.85 default remains conservative (it hands off more, never less).

Routing quality (AUROC)

AUROC measures how well the probe separates wrong answers from right ones (higher = better, 0.5 is random, 1.0 is perfect):

Hold-out Modality Cactus Hybrid Token Entropy
MMLU text MCQ 0.770 0.697
MMLU-Pro text MCQ 0.771 0.692
ARC-Easy text MCQ 0.888 0.655
ARC-Challenge text MCQ 0.834 0.646
GSM8K (3-shot) text gen 0.782 0.731
MMBench-EN-Dev vision MCQ 0.840 0.435
ChartQA vision QA 0.779 0.615
DocVQA vision QA 0.781 0.512
MMAU audio MCQ 0.789 0.517
GigaSpeech audio 0.876 0.343
Earnings-22 audio 0.839 0.323
LibriSpeech audio 0.822 0.427
Mean 0.814 0.549

The strongest result: the probe was trained on zero audio data, yet achieves 0.79–0.88 AUROC on four audio benchmarks (two transcription, one audio MCQ, one out-of-domain transcription). This rules out surface-level explanations: the probe is reading a modality-independent correctness signal from the hidden state, not memorizing patterns from training data.

All formats

All Cactus Hybrid builds live in the Cactus Hybrid collection: Transformers · GGUF / llama.cpp · MLX · Cactus engine. Copy-paste quickstarts for every engine: github.com/cactus-compute/cactus-hybrid.

License

Gemma is provided under and subject to the Gemma Terms of Use. This derivative includes the Cactus handoff probe head.

Downloads last month
211
GGUF
Model size
5B params
Architecture
gemma-4-e2b-it-hybrid
Hardware compatibility
Log In to add your hardware

4-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Cactus-Compute/gemma-4-e2b-it-hybrid-GGUF

Quantized
(2)
this model

Collection including Cactus-Compute/gemma-4-e2b-it-hybrid-GGUF