Instructions to use Mittai17/gemma4-e2b-medtrain-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Mittai17/gemma4-e2b-medtrain-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M
Use Docker
docker model run hf.co/Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use Mittai17/gemma4-e2b-medtrain-gguf with Ollama:
ollama run hf.co/Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M
- Unsloth Desktop
- Pi
How to use Mittai17/gemma4-e2b-medtrain-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Mittai17/gemma4-e2b-medtrain-gguf with Docker Model Runner:
docker model run hf.co/Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M
- Lemonade
How to use Mittai17/gemma4-e2b-medtrain-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M
Run and chat with the model
lemonade run user.gemma4-e2b-medtrain-gguf-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Mittai17/gemma4-e2b-medtrain-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Mittai17/gemma4-e2b-medtrain-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Mittai17/gemma4-e2b-medtrain-gguf:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Gemma 4 E2B-it — MedTrain GGUF (Q4_K_M)
A GGUF quantization of Google Gemma 4 E2B-it LoRA-fine-tuned as a health screening and first aid assistant for India. The LoRA adapter (r=8, alpha=32) was merged into the base model, converted to GGUF (bf16), and quantized to Q4_K_M for on-device use on a ~4 GB RAM phone.
- Base model:
google/gemma-4-E2B-it - Quantization:
Q4_K_M(≈2.9 GB) - Architecture: Gemma4 (per-layer token embeddings, per-layer scalars, hybrid sliding-window / full attention)
Files
| File | Size | Purpose |
|---|---|---|
gemma4-e2b-medtrain-Q4_K_M.gguf |
2.9 GB | Deployment model |
Behavior
The model follows the trained response structure:
- PRECAUTIONS — what to do to avoid making things worse and what NOT to do
- SOLUTIONS — clear, step-by-step actions to take right now
- WHEN TO GET HELP — which signs require urgent or emergency care
- Indian emergency numbers: 112 (national), 108 (ambulance), 102/104 (health helplines), 100 (police), 101 (fire)
It responds in the same language the user writes in (English, Hindi, Tamil, Telugu, Bengali, Kannada, Marathi, Malayalam, Gujarati, Punjabi, Urdu, Odia, Assamese, and other Indian languages).
This is a "thinking" model — it emits [Start thinking] / [End thinking] delimiters. On-device apps should render or trim these.
System prompt
Use this as the system prompt:
You are a health screening and first aid assistant for India. Respond in the SAME language the user writes in: English, Hindi, Tamil, Telugu, Bengali, Kannada, Marathi, Malayalam, Gujarati, Punjabi, Urdu, Odia, Assamese, or other Indian languages. Keep answers clear and simple. Always explain the PRECAUTIONS (what to do to avoid making things worse and what NOT to do), the SOLUTIONS (clear step-by-step actions to take right now), and WHEN TO GET HELP (which signs require urgent or emergency care). For emergencies in India: call 112 (national emergency, works in all states), 108 (ambulance), 102 / 104 (health helplines), 100 (police), 101 (fire). Emergency numbers may vary by state and operator. Always prioritize urgent medical attention for emergency symptoms. Use cautious, uncertain language and never give a diagnosis.
Usage with llama.cpp
llama-cli -m gemma4-e2b-medtrain-Q4_K_M.gguf \
--threads 8 --ctx-size 2048 \
-p "USER: A person got an electric shock. What should we do?\nASSISTANT:" \
-n 1200 -st
Disclaimer
This model is a basic health-screening assistant, not a clinical diagnostic tool, and must not be used as a substitute for professional medical advice. All outputs are preliminary and require independent verification.
- Downloads last month
- -
4-bit