Merv on Gemma 4 E4B

Gemma 4 E4B (7.5B parameters), fine-tuned into Merv, a two-headed robot: every answer comes from two personalities at once, each wrapped in its own tag so a chat interface can split the voices.

  • Mervin: gloomy and sardonic.
  • Mervis: relentlessly cheerful.
User:    What is 2+2?
Mervin:  A trivial sum, naturally assigned to me because apparently no one else in the universe can survive counting to four.
Mervis:  Marvelous! That answer practically sparkles with useful little possibilities, like a sunrise wearing sensible shoes.

Merv is a demonstration of teaching a small model a precise style and output format from a few hundred examples, then running it entirely on the user's own device (the Phi-4-mini version runs in a web browser through WebGPU). The training data (about 260 examples) and the build notebook are at github.com/freeideas/merv. Use follows the base model's licence.

Run it

Files: model-q4_k_m.gguf, 4-bit (Q4_K_M).

Run it with llama.cpp:

llama-cli -hf freeideas/merv-gemma4e4b -cnv

or load the GGUF file in LM Studio, Ollama, or llama-cpp-python.

About the author

Made by Carl Free, an AI engineer who trains small, specialized models that beat large general models on narrow tasks at a fraction of the cost. Résumé, demos, and contact: ordinarydata.com/resume. More models: huggingface.co/freeideas.

Downloads last month
32
GGUF
Model size
8B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support