Configuration Parsing Warning:In config.json: "num_experts" must be a number

Cactus Hybrid — Gemma 4 E2B (MLX, 4-bit)

A small, on-device model is fast and private, but sometimes wrong. At Cactus we post-train models to know when they are wrong: we ship probes inside the checkpoint that score every answer with a confidence between 0 and 1, returned as structured data (never parsed out of the answer text). Answer on-device when confidence is high; re-route to a bigger model when it's low:

if confidence < 0.85:
    answer = ask_a_bigger_model(prompt)

This repo holds the MLX-converted 4-bit build of Cactus-Compute/gemma-4-e2b-it-hybrid. The architecture ships in this repo via mlx-lm's model_file remote-code mechanism (mlx-lm ≥ 0.30.1); the probe head is stored float32, never quantized (only the trunk is 4-bit, g64).

Benchmarks

Gemma 4 E2B Hybrid, the smallest Gemma model, matches Gemini 3.1 Flash-Lite on most benchmarks by routing only 15–35% of queries to Flash-Lite and running the rest itself:

Benchmark Handoff to match Flash-Lite (FP16) At 4-bit At 3-bit
ChartQA 15–20% 25–30% 40–50%
MMBench 30–35% 40–45% 50–55%
LibriSpeech 25–30% 35–40% 55–65%
GigaSpeech 30–35% 40–45% 50–55%
MMAU 30–35% 35–40% 50–55%
MMLU-Pro 45–55% ~90% n/a

Quantisation quality is measured on Cactus Quants, which performs well at uniform quantization; developers are encouraged to benchmark Unsloth, GGUF, and MLX quantization independently.

Quickstart

# pip install mlx-lm
import re
from mlx_lm import load, generate

model, tokenizer = load(
    "Cactus-Compute/gemma-4-e2b-it-hybrid-mlx",
    tokenizer_config={"trust_remote_code": True},
)

messages = [{"role": "user", "content": "What is the capital of France?"}]
answer = generate(
    model,
    tokenizer,
    prompt=tokenizer.apply_chat_template(messages, add_generation_prompt=True),
    max_tokens=512,
)
# the checkpoint reasons before answering; keep only the final answer
answer = re.split(r"<\|?channel\|?>", answer)[-1]
answer = re.sub(r"^(thought|final)\b\s*", "", answer).strip()
print(answer)
print("confidence:", model.last_confidence)

Confidence on MLX is exposed through the Python API — model.last_confidence after generation (or model.confidence(num_tokens=N)). mlx_lm.server serves the model fine but cannot add a confidence field to its responses, so read the score in-process.

Calibration notes

  • On matched generation trajectories the 4-bit probe drift vs the bf16 reference is under 0.01.
  • The 4-bit trunk can shift the greedy thinking/non-thinking boundary versus bf16: some prompts enter the thinking channel where bf16 answers directly, and the probe legitimately scores those different generations lower. Easy vs hard ordering is fully preserved.

Routing quality (AUROC)

AUROC measures how well the probe separates wrong answers from right ones (higher = better, 0.5 is random, 1.0 is perfect):

Hold-out Modality Cactus Hybrid Token Entropy
MMLU text MCQ 0.770 0.697
MMLU-Pro text MCQ 0.771 0.692
ARC-Easy text MCQ 0.888 0.655
ARC-Challenge text MCQ 0.834 0.646
GSM8K (3-shot) text gen 0.782 0.731
MMBench-EN-Dev vision MCQ 0.840 0.435
ChartQA vision QA 0.779 0.615
DocVQA vision QA 0.781 0.512
MMAU audio MCQ 0.789 0.517
GigaSpeech audio 0.876 0.343
Earnings-22 audio 0.839 0.323
LibriSpeech audio 0.822 0.427
Mean 0.814 0.549

The strongest result: the probe was trained on zero audio data, yet achieves 0.79–0.88 AUROC on four audio benchmarks (two transcription, one audio MCQ, one out-of-domain transcription). This rules out surface-level explanations: the probe is reading a modality-independent correctness signal from the hidden state, not memorizing patterns from training data.

All formats

All Cactus Hybrid builds live in the Cactus Hybrid collection: Transformers · GGUF / llama.cpp · MLX · Cactus engine. Copy-paste quickstarts for every engine: github.com/cactus-compute/cactus-hybrid.

License

Gemma is provided under and subject to the Gemma Terms of Use. This derivative includes the Cactus handoff probe head.

Downloads last month
60
Safetensors
Model size
0.7B params
Tensor type
F32
·
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Cactus-Compute/gemma-4-e2b-it-hybrid-mlx

Quantized
(2)
this model

Collection including Cactus-Compute/gemma-4-e2b-it-hybrid-mlx