Instructions to use Kenntkim66/snowclaw-gemma4-e2b-ft-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Kenntkim66/snowclaw-gemma4-e2b-ft-gguf with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Kenntkim66/snowclaw-gemma4-e2b-ft-gguf") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Kenntkim66/snowclaw-gemma4-e2b-ft-gguf", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Kenntkim66/snowclaw-gemma4-e2b-ft-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16 # Run inference directly in the terminal: llama cli -hf Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16 # Run inference directly in the terminal: llama cli -hf Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16 # Run inference directly in the terminal: ./llama-cli -hf Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16
Use Docker
docker model run hf.co/Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16
- LM Studio
- Jan
- vLLM
How to use Kenntkim66/snowclaw-gemma4-e2b-ft-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Kenntkim66/snowclaw-gemma4-e2b-ft-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kenntkim66/snowclaw-gemma4-e2b-ft-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16
- SGLang
How to use Kenntkim66/snowclaw-gemma4-e2b-ft-gguf with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Kenntkim66/snowclaw-gemma4-e2b-ft-gguf" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kenntkim66/snowclaw-gemma4-e2b-ft-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Kenntkim66/snowclaw-gemma4-e2b-ft-gguf" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kenntkim66/snowclaw-gemma4-e2b-ft-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use Kenntkim66/snowclaw-gemma4-e2b-ft-gguf with Ollama:
ollama run hf.co/Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16
- Unsloth Studio
How to use Kenntkim66/snowclaw-gemma4-e2b-ft-gguf with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Kenntkim66/snowclaw-gemma4-e2b-ft-gguf to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Kenntkim66/snowclaw-gemma4-e2b-ft-gguf to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Kenntkim66/snowclaw-gemma4-e2b-ft-gguf to start chatting
- Pi
How to use Kenntkim66/snowclaw-gemma4-e2b-ft-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use Kenntkim66/snowclaw-gemma4-e2b-ft-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use Kenntkim66/snowclaw-gemma4-e2b-ft-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use Kenntkim66/snowclaw-gemma4-e2b-ft-gguf with Docker Model Runner:
docker model run hf.co/Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16
- Lemonade
How to use Kenntkim66/snowclaw-gemma4-e2b-ft-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Kenntkim66/snowclaw-gemma4-e2b-ft-gguf:BF16
Run and chat with the model
lemonade run user.snowclaw-gemma4-e2b-ft-gguf-BF16
List all available models
lemonade list
SnowClaw — Fine-tuned Gemma 4 E2B for Privacy-First Tool Use
A fine-tuned Gemma 4 E2B model optimized for on-device AI tool use in the SnowClaw desktop agent. Achieves 100% tool use accuracy (7/7) on our evaluation set.
Model Details
| Property | Value |
|---|---|
| Base Model | google/gemma-4-E2B-it (Gemma 4 E2B Instruct) |
| Method | LoRA (rank=64, alpha=64) via Unsloth |
| Training Data | 2,000 synthetic tool-use examples |
| Epochs | 15 |
| Final Loss | 0.040 |
| Hardware | NVIDIA RTX 3090 (24GB VRAM) |
| Quantization | Q4_K_M (GGUF) |
| Vision | Multimodal — includes mmproj for image understanding |
Files
| File | Size | Description |
|---|---|---|
gemma-4-e2b-it.Q4_K_M.gguf |
3.2 GB | Main model (Q4_K_M quantized) |
gemma-4-e2b-it.BF16-mmproj.gguf |
942 MB | Vision projector (multimodal image encoder) |
Modelfile |
205 B | Ollama registration file |
Intended Use
SnowClaw is a privacy-first desktop AI agent that runs entirely on-device. This model is fine-tuned for:
- Tool Use: Executing system commands, browsing files, managing contacts/calendar
- Code Generation: Writing and executing Python/AppleScript in a sandboxed environment
- Screenshot Analysis: Understanding screen content via vision capabilities
- Privacy: All processing stays local — zero data leaves the device
How to Use
With Ollama
# Download the GGUF files, then register with Ollama:
ollama create snowclaw -f Modelfile
# Run
ollama run snowclaw
With llama.cpp
# Text only
./llama-cli -m gemma-4-e2b-it.Q4_K_M.gguf -p "List files in my Downloads folder"
# With vision (multimodal)
./llama-mtmd-cli \
-m gemma-4-e2b-it.Q4_K_M.gguf \
--mmproj gemma-4-e2b-it.BF16-mmproj.gguf
Training Details
LoRA Configuration
{
"r": 64,
"lora_alpha": 64,
"target_modules": ["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"],
"lora_dropout": 0,
"task_type": "CAUSAL_LM"
}
Training Curve
| Step | Loss |
|---|---|
| 10 | 8.147 |
| 20 | 1.074 |
| 30 | 0.469 |
| 100 | 0.099 |
| 500 | 0.043 |
| 1000 | 0.038 |
| 1875 | 0.040 |
Dataset
2,000 synthetic examples covering:
- System tool invocations (file management, process control)
- Contact and calendar queries
- Device information retrieval
- Multi-step task planning
- Safety-aware refusals
Evaluation
| Metric | Score |
|---|---|
| Tool Use Accuracy | 7/7 (100%) |
| Correct Tool Selection | 7/7 |
| Parameter Extraction | 7/7 |
Part of SnowClaw
SnowClaw is a privacy-first AI agent built for the Google Gemma Hackathon. It features:
- On-device inference via bundled Ollama
- Dual security modes: Paranoid (fully offline) / Smart Search (local + anonymous SearXNG)
- E2E encrypted communication between desktop and mobile
- Hardware-aware model selection (auto-detects CPU/GPU/RAM)
Author
Kennt Kim — Calida Lab
License
Apache 2.0 (following Gemma's license terms)
- Downloads last month
- 13
4-bit
Model tree for Kenntkim66/snowclaw-gemma4-e2b-ft-gguf
Evaluation results
- Tool Use (7/7)self-reported100.000