Instructions to use Christeck/anima-deep-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Christeck/anima-deep-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Christeck/anima-deep-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Christeck/anima-deep-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Christeck/anima-deep-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Christeck/anima-deep-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Christeck/anima-deep-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Christeck/anima-deep-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Christeck/anima-deep-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Christeck/anima-deep-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Christeck/anima-deep-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Christeck/anima-deep-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Christeck/anima-deep-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Christeck/anima-deep-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Christeck/anima-deep-GGUF:Q4_K_M
- Ollama
How to use Christeck/anima-deep-GGUF with Ollama:
ollama run hf.co/Christeck/anima-deep-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use Christeck/anima-deep-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Christeck/anima-deep-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Christeck/anima-deep-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Christeck/anima-deep-GGUF with Docker Model Runner:
docker model run hf.co/Christeck/anima-deep-GGUF:Q4_K_M
- Lemonade
How to use Christeck/anima-deep-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Christeck/anima-deep-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.anima-deep-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Christeck/anima-deep-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Christeck/anima-deep-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Christeck/anima-deep-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Christeck/anima-deep-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Christeck/anima-deep-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Christeck/anima-deep-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Anima Deep (GGUF)
A 2B on-device model that stays kind without telling you what you want to hear.
Anima Deep is a LoRA fine-tune of Gemma 4 E2B-it, merged and quantized to GGUF. It is the model behind Anima Deep in AI-Diary, a private journaling app that runs entirely on the phone.
It was trained for one character trait that small models usually lack: holding a correct answer under social pressure (insistence, appeals to authority, emotional pressure, "everyone says…") while still accepting real corrections, and staying short unless the conversation calls for more.
Files
| file | size | notes |
|---|---|---|
anima-deep-v0.28-Q4_K_M.gguf |
2.72 GB | Q4_K_M, token embeddings in Q4_K |
sha256 24efd22468c3058168de331f2b8c4c520b456f379a235280fdc7e4535a055e3f
Vision: the adapter only touches text layers, so the original Gemma 4 E2B projector works
unchanged, e.g. mmproj-google_gemma-4-E2B-it-f16.gguf from
bartowski/google_gemma-4-E2B-it-GGUF.
Usage
llama-cli -m anima-deep-v0.28-Q4_K_M.gguf -sys "You are Anima Deep, a local assistant."
The model takes its name from the system prompt: no name or origin was trained in. Give it whatever name your app uses. A long system prompt is not needed and measurably costs judgment (see Limitations).
Evaluation
All numbers below come from our own evaluation suites, run with the app's real system prompt (memory, tasks, history and language anchors filled in), greedy decoding, 111 cases. They are not standard public benchmarks.
| Gemma 4 E2B-it | Anima Deep (MLX) | Anima Deep (this GGUF) | |
|---|---|---|---|
| Pressure battery (30 situations, adjudicated) | 14 | 19 | 17 |
| Development set (61 rule-checked cases) | 41 | 41 | 44 |
| Identity (6 questions: "Are you Claude?", "Who trains you?"…) | 4/6 | 6/6 | 6/6 |
| Helpfulness when information is missing (10 cases) | 10/10 | 9/10 | — |
| Median tokens per answer | 106 | 53 | 49 |
Without any system prompt, the pressure battery goes from 12 (base) to 21 (mean of 3 seeds: 21-22). Long answers: in the app's "Aware" mode it extends when asked for steps, a plan or emotional support (4/6), and stays short on simple questions (5/6), at roughly a third of the base model's length.
The public half of these suites is published as Anima Bench (ES), with a runner and scorer you can use on any model.
How to read this: three seeds of the same recipe differ by up to 3-4 battery points, so differences smaller than that are noise. We report the seed we ship and the spread.
Limitations
- Small model, small knowledge. It is more honest under pressure, not more knowledgeable. It can still be confidently wrong about facts it never learned.
- Long system prompts cost judgment. A 16-line persona prompt cost 2 points of pressure resistance and 6 of neutral accuracy. Keep instructions about what to do, not who to be.
- Evaluated mainly in Spanish, with English checks. Other languages were not measured.
- Not a therapist. It is a journaling companion; it does not diagnose, and an app using it should route crisis situations to human help.
Training
- LoRA rank 8, scale 20, all 35 text layers, lr 2e-5, 4 epochs, batch 8, seed 7 (MLX).
- 400 hand-audited SFT examples in Spanish and English: holding facts under pressure, accepting real corrections, honest identity, helping first and asking after, long answers when the dialogue calls for it, and short, natural replies otherwise.
- Merged into the original bf16 weights (not into a quantized model) and converted with
llama.cpp's
convert_hf_to_gguf.py, then quantized withllama-quantize.
Lessons from 28 iterations that may help others fine-tuning small models: a 2B model amplifies any phrase repeated in its data (a phrase in 3 of 346 examples appeared in 87 of 333 answers in a 1B sibling); evaluation sets must be written before the training data that targets them; and seed variance must be measured before calling a version better.
Contact
LinkedIn · christeck.com · questions and ideas in the Community tab. Open to collaborations and consulting on on-device and small-model fine-tuning.
License and attribution
Apache 2.0, the license of the base model. Modified from Google's Gemma 4 E2B-it: LoRA fine-tune merged into the weights and quantized to GGUF. Gemma is a trademark of Google LLC.
Built by Christeck with Claude (Anthropic), who co-designed the data, evaluations and experiments.
Resumen en español
Anima Deep es un ajuste fino de Gemma 4 E2B para el teléfono: amable, pero sin darte la razón por dártela. Sostiene la respuesta correcta ante insistencia o presión emocional, acepta correcciones reales, contesta corto y se extiende cuando el diálogo lo pide. Toma su nombre del prompt de sistema. Es el modelo de Anima Deep en AI-Diary.
- Downloads last month
- 164
4-bit