Instructions to use KiwiMate/KiwiMate-Mini-Preview with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use KiwiMate/KiwiMate-Mini-Preview with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf KiwiMate/KiwiMate-Mini-Preview:Q4_K_M # Run inference directly in the terminal: llama cli -hf KiwiMate/KiwiMate-Mini-Preview:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf KiwiMate/KiwiMate-Mini-Preview:Q4_K_M # Run inference directly in the terminal: llama cli -hf KiwiMate/KiwiMate-Mini-Preview:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf KiwiMate/KiwiMate-Mini-Preview:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf KiwiMate/KiwiMate-Mini-Preview:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf KiwiMate/KiwiMate-Mini-Preview:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf KiwiMate/KiwiMate-Mini-Preview:Q4_K_M
Use Docker
docker model run hf.co/KiwiMate/KiwiMate-Mini-Preview:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use KiwiMate/KiwiMate-Mini-Preview with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "KiwiMate/KiwiMate-Mini-Preview" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "KiwiMate/KiwiMate-Mini-Preview", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/KiwiMate/KiwiMate-Mini-Preview:Q4_K_M
- Ollama
How to use KiwiMate/KiwiMate-Mini-Preview with Ollama:
ollama run hf.co/KiwiMate/KiwiMate-Mini-Preview:Q4_K_M
- Unsloth Studio
How to use KiwiMate/KiwiMate-Mini-Preview with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for KiwiMate/KiwiMate-Mini-Preview to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for KiwiMate/KiwiMate-Mini-Preview to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for KiwiMate/KiwiMate-Mini-Preview to start chatting
- Pi
How to use KiwiMate/KiwiMate-Mini-Preview with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf KiwiMate/KiwiMate-Mini-Preview:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "KiwiMate/KiwiMate-Mini-Preview:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use KiwiMate/KiwiMate-Mini-Preview with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf KiwiMate/KiwiMate-Mini-Preview:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "KiwiMate/KiwiMate-Mini-Preview:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use KiwiMate/KiwiMate-Mini-Preview with Docker Model Runner:
docker model run hf.co/KiwiMate/KiwiMate-Mini-Preview:Q4_K_M
- Lemonade
How to use KiwiMate/KiwiMate-Mini-Preview with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull KiwiMate/KiwiMate-Mini-Preview:Q4_K_M
Run and chat with the model
lemonade run user.KiwiMate-Mini-Preview-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use KiwiMate/KiwiMate-Mini-Preview with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf KiwiMate/KiwiMate-Mini-Preview:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default KiwiMate/KiwiMate-Mini-Preview:Q4_K_M
Run Hermes
hermes
- Atomic Chat
KiwiMate-Mini-Preview
KiwiMate-Mini-Preview is a lightweight, New Zealand–flavoured conversational language model, fine-tuned for the KiwiMate AI companion app. It is the smallest model in the KiwiMate model family and is designed for fast, low-cost inference on the app's free and lower-tier subscription plans.
⚠️ Preview status: This is a prototype release. The name reflects its preview status — expect breaking changes, retraining, and behavioural shifts before a stable v1 release.
Model Details
| Developed by | KiwiMate / KyleCodeKiwi |
| Base model | Llama 3.2 3B Instruct |
| Architecture | Llama |
| Parameters | ~3.21B |
| Fine-tuning framework | Unsloth |
| License | Apache 2.0 |
| Languages | English (with New Zealand English and Te Reo Māori vocabulary coverage) |
| Model class | AutoModelForCausalLM |
Intended Use
KiwiMate-Mini-Preview is intended as the default conversational backend for the KiwiMate app, providing:
- General-purpose chat and assistant-style conversation
- New Zealand cultural and "Kiwi" context awareness (slang, geography, fun facts)
- Light Te Reo Māori vocabulary recognition and use
- Lore and knowledge specific to the KiwiMate app itself ("KiwiMate Origin" data)
- Lightweight knowledge support for in-app mini-games
It is not intended for high-stakes, medical, legal, or financial advice, and should not be relied on as an authoritative source on Māori language or tikanga — for genuinely sensitive Te Reo or cultural content, defer to community-governed resources.
Training Data
Fine-tuned on the KiwiMate/KiwiMate-Mini-training dataset, organised into categories including:
- NZ English
- Te Reo Māori
- KiwiMate Origin (app-specific lore/identity)
- NZ Fun Facts
- MiniGame Knowledge
Files & Quantizations
Distributed as safetensors (full precision) and GGUF quantizations for efficient local/edge inference:
| Format | Use case |
|---|---|
| F16 | Highest fidelity, largest size |
| Q6_K | Near-lossless, smaller footprint |
| Q4_K_M | Balanced quality/size — recommended default for on-device use |
| Q2_K_L | Smallest footprint, lowest fidelity |
Deployment
Served in production via a Hugging Face Inference Endpoint on a T4 GPU with scale-to-zero, fronted by a Supabase Edge Function (OpenAI-compatible proxy) that routes KiwiMate app traffic to this and other KiwiMate model endpoints behind a single API.
Known Limitations
- A server-side mitigation is in place for an occasional role-bleed / over-generation issue (the model sometimes continuing past
<|eot_id|>), handled via stop-sequence aliases and trimming at the proxy layer. - The long-term fix — adding
<|eot_id|>(token ID 128009) properly to the training loss andgeneration_config.json— is planned for a future retraining pass rather than this preview. - As a 3B-parameter model, reasoning depth and factual reliability are limited compared to larger models; it is tuned for speed and personality over raw capability.
License
Released under the Apache 2.0 license, consistent with the open weights commitment for the KiwiMate model family.
- Downloads last month
- 109
Model tree for KiwiMate/KiwiMate-Mini-Preview
Base model
meta-llama/Llama-3.2-3B-Instruct