Anima Deep (GGUF)

A 2B on-device model that stays kind without telling you what you want to hear.

Anima Deep is a LoRA fine-tune of Gemma 4 E2B-it, merged and quantized to GGUF. It is the model behind Anima Deep in AI-Diary, a private journaling app that runs entirely on the phone.

It was trained for one character trait that small models usually lack: holding a correct answer under social pressure (insistence, appeals to authority, emotional pressure, "everyone says…") while still accepting real corrections, and staying short unless the conversation calls for more.

Files

file size notes
anima-deep-v0.28-Q4_K_M.gguf 2.72 GB Q4_K_M, token embeddings in Q4_K

sha256 24efd22468c3058168de331f2b8c4c520b456f379a235280fdc7e4535a055e3f

Vision: the adapter only touches text layers, so the original Gemma 4 E2B projector works unchanged, e.g. mmproj-google_gemma-4-E2B-it-f16.gguf from bartowski/google_gemma-4-E2B-it-GGUF.

Usage

llama-cli -m anima-deep-v0.28-Q4_K_M.gguf -sys "You are Anima Deep, a local assistant."

The model takes its name from the system prompt: no name or origin was trained in. Give it whatever name your app uses. A long system prompt is not needed and measurably costs judgment (see Limitations).

Evaluation

All numbers below come from our own evaluation suites, run with the app's real system prompt (memory, tasks, history and language anchors filled in), greedy decoding, 111 cases. They are not standard public benchmarks.

Gemma 4 E2B-it Anima Deep (MLX) Anima Deep (this GGUF)
Pressure battery (30 situations, adjudicated) 14 19 17
Development set (61 rule-checked cases) 41 41 44
Identity (6 questions: "Are you Claude?", "Who trains you?"…) 4/6 6/6 6/6
Helpfulness when information is missing (10 cases) 10/10 9/10 —
Median tokens per answer 106 53 49

Without any system prompt, the pressure battery goes from 12 (base) to 21 (mean of 3 seeds: 21-22). Long answers: in the app's "Aware" mode it extends when asked for steps, a plan or emotional support (4/6), and stays short on simple questions (5/6), at roughly a third of the base model's length.

The public half of these suites is published as Anima Bench (ES), with a runner and scorer you can use on any model.

How to read this: three seeds of the same recipe differ by up to 3-4 battery points, so differences smaller than that are noise. We report the seed we ship and the spread.

Limitations

  • Small model, small knowledge. It is more honest under pressure, not more knowledgeable. It can still be confidently wrong about facts it never learned.
  • Long system prompts cost judgment. A 16-line persona prompt cost 2 points of pressure resistance and 6 of neutral accuracy. Keep instructions about what to do, not who to be.
  • Evaluated mainly in Spanish, with English checks. Other languages were not measured.
  • Not a therapist. It is a journaling companion; it does not diagnose, and an app using it should route crisis situations to human help.

Training

  • LoRA rank 8, scale 20, all 35 text layers, lr 2e-5, 4 epochs, batch 8, seed 7 (MLX).
  • 400 hand-audited SFT examples in Spanish and English: holding facts under pressure, accepting real corrections, honest identity, helping first and asking after, long answers when the dialogue calls for it, and short, natural replies otherwise.
  • Merged into the original bf16 weights (not into a quantized model) and converted with llama.cpp's convert_hf_to_gguf.py, then quantized with llama-quantize.

Lessons from 28 iterations that may help others fine-tuning small models: a 2B model amplifies any phrase repeated in its data (a phrase in 3 of 346 examples appeared in 87 of 333 answers in a 1B sibling); evaluation sets must be written before the training data that targets them; and seed variance must be measured before calling a version better.

Contact

LinkedIn · christeck.com · questions and ideas in the Community tab. Open to collaborations and consulting on on-device and small-model fine-tuning.

License and attribution

Apache 2.0, the license of the base model. Modified from Google's Gemma 4 E2B-it: LoRA fine-tune merged into the weights and quantized to GGUF. Gemma is a trademark of Google LLC.

Built by Christeck with Claude (Anthropic), who co-designed the data, evaluations and experiments.


Resumen en español

Anima Deep es un ajuste fino de Gemma 4 E2B para el teléfono: amable, pero sin darte la razón por dártela. Sostiene la respuesta correcta ante insistencia o presión emocional, acepta correcciones reales, contesta corto y se extiende cuando el diálogo lo pide. Toma su nombre del prompt de sistema. Es el modelo de Anima Deep en AI-Diary.

Downloads last month
164
GGUF
Model size
5B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Christeck/anima-deep-GGUF

Finetuned
(387)
this model