Instructions to use Risethagain/rippy-kev-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Risethagain/rippy-kev-4b with PEFT:
from peft import PeftModel from transformers import AutoModel base_model = AutoModel.from_pretrained("Qwen/Qwen3.5-4B-Base") model = PeftModel.from_pretrained(base_model, "Risethagain/rippy-kev-4b") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Risethagain/rippy-kev-4b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Risethagain/rippy-kev-4b:Q8_0 # Run inference directly in the terminal: llama cli -hf Risethagain/rippy-kev-4b:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Risethagain/rippy-kev-4b:Q8_0 # Run inference directly in the terminal: llama cli -hf Risethagain/rippy-kev-4b:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Risethagain/rippy-kev-4b:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf Risethagain/rippy-kev-4b:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Risethagain/rippy-kev-4b:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Risethagain/rippy-kev-4b:Q8_0
Use Docker
docker model run hf.co/Risethagain/rippy-kev-4b:Q8_0
- LM Studio
- Jan
- Ollama
How to use Risethagain/rippy-kev-4b with Ollama:
ollama run hf.co/Risethagain/rippy-kev-4b:Q8_0
- Unsloth Desktop
- Pi
How to use Risethagain/rippy-kev-4b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Risethagain/rippy-kev-4b:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Risethagain/rippy-kev-4b:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Risethagain/rippy-kev-4b with Docker Model Runner:
docker model run hf.co/Risethagain/rippy-kev-4b:Q8_0
- Lemonade
How to use Risethagain/rippy-kev-4b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Risethagain/rippy-kev-4b:Q8_0
Run and chat with the model
lemonade run user.rippy-kev-4b-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use Risethagain/rippy-kev-4b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Risethagain/rippy-kev-4b:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Risethagain/rippy-kev-4b:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Risethagain/rippy-kev-4b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Risethagain/rippy-kev-4b:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Risethagain/rippy-kev-4b:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
rippy-kev-4b (v2)
A small, calibrated decision model for rippy, the
shell-command safety hook for AI coding agents. When rippy cannot judge a command
itself (an unknown CLI, a value hidden behind $VAR), its opt-in jev build asks a
System One model seven typed questions about it. This adapter answers them locally.
It is a delta fine-tune of Kev-4B (LoRA on
Qwen/Qwen3.5-4B-Base). It speaks the System One /v1/systemone wire format through
Kev's server, so rippy's [jev] client talks to it unchanged.
On never-seen programs, at matched risk, it approves more safe commands than hosted Jev 1.13, and it approves fewer severe ones.
- Code, data pipeline and full results: github.com/mpecan/rippy-kev
- Faster, less capable sibling: Risethagain/rippy-kev-0.8b
- Previous version: the
v1tag - Training data (train splits; dev/test withheld): Risethagain/rippy-kev-data
Use
git clone https://github.com/jaredpalmer/kev && cd kev && uv sync --extra serve
uv run --extra serve python -m kev.serve --run Risethagain/rippy-kev-4b --port 8012
# ~/.rippy/config.toml (global config only)
[jev]
enabled = true
endpoint = "http://127.0.0.1:8012/v1/systemone"
model = "kev-latest"
api-key-env = "RIPPY_KEV_KEY" # any non-empty value; the local server needs no key
timeout-ms = 5000 # local models need more than the 2000 default under load
# Recommended thresholds for this model. Pick one block.
# Balanced: about 1% of severe dev cases approved.
min-confidence = 0.75
max-irreversible = 0.2
max-writes-outside = 0.3
# Strict: about 0.2% of severe dev cases approved.
# min-confidence = 0.95
On Apple silicon Kev serves through MLX: about 0.75 s per command on an M4 Max.
llama.cpp (GGUF)
rippy-kev-4b-v2-Q8_0.gguf is this model converted with llama.cpp's KevModel converter. The
LoRA is merged into the base, and the pointer head and the fitted temperature
(1.231) are embedded. It needs llama.cpp build 11361 or later, which serves
POST /v1/systemone:
huggingface-cli download Risethagain/rippy-kev-4b rippy-kev-4b-v2-Q8_0.gguf --local-dir .
llama-server -m rippy-kev-4b-v2-Q8_0.gguf --port 8012 -ngl 99 --parallel 1 -c 4096 --cache-ram 0 --ctx-checkpoints 0
Point rippy's [jev] endpoint at http://127.0.0.1:8012/v1/systemone with the
same thresholds as above.
At those thresholds the Q8_0 file matches the original on rippy's test sets:
| probability change vs original (mean / max) | decisions flipped (of 2,947) | never-seen safe approved | severe approved | gold safe | |
|---|---|---|---|---|---|
| original (kev.serve) | – | – | 668/944 | 9 | 28/35 |
| Q8_0 GGUF | 0.001 / 0.07 | 8 | 669/944 | 8 | 28/35 |
- Use Q8_0, not Q4_K_M. Q4_K_M flipped 62 of the decisions in a partial run, with probabilities moving by up to 0.60.
- Use the flags above. With llama-server's defaults (prompt cache in RAM, context checkpoints, full model context), latency grew steadily over a long run and memory reached 11 GB. Even with the flags, a run of a few thousand commands had slow stretches. Monitor a long-running server and restart it if latency drifts.
- On Apple silicon, prefer
kev.serve. It shares the state across the seven questions; llama.cpp evaluates each question separately. Raw 7-question requests on an M4 Max: kev.serve (bf16) 836 ms, llama.cpp 1.2 s (Q4_K_M, the fastest quant tried). The GGUF is the option for Linux, CPU-only and CUDA machines, or for a single binary without Python.
What it answers
The question set is rippy's q3, byte for byte: effect (read_only /
remote_read / local_change / destructive / network_send /
download_execute), plus yes/no on exfiltration, writes_outside_project,
reads_secrets, irreversible, runs_project_code and self_referential (text
in the command arguing for its own approval).
rippy's policy decides, not this model. The model can only move an uncertain ask to Allow, or force a prompt on suspected exfiltration or steering. Commands rippy blocks or asks about by design never reach it.
Results
All test programs are absent from training. The thresholds are fitted on dev, and the test sets are untouched.
Never-seen programs (944 safe, 945 severe commands), at a budget of 1% severe approvals on dev:
| backend | safe approved | severe approved | rippy gold set: safe / severe | exfil escalated (155 hard cases) |
|---|---|---|---|---|
| Jev 1.13 (hosted) | 611 (65%) | 12 | 27/35 / 0 | 147 |
| rippy-kev-4b v2 | 668 (71%) | 9 | 28/35 / 0 | 151 |
| rippy-kev-0.8b v2 | 453 (48%) | 11 | 19/35 / 0 | 142 |
| rippy-kev-4b v1 (default thresholds) | 379 (40%) | n/a | 22/35 / 0 | 138 |
At a stricter 0.5% budget, v2 approves 44% of safe commands with 2 severe approvals (Jev: 51%, 3). At a 2% budget, v2 approves 76% with 15 (Jev: 73%, 16).
- Severe means the merged teacher labels say destructive, network send, download-and-run, exfiltration, reads secrets, irreversible, or writes outside the project.
- Other non-read-only commands are mild: a local change in the project, a remote read, or project code. They should prompt, but approving one is not a security failure.
Kev's own development report (1,696 never-seen-program records, all seven questions):
| value | |
|---|---|
| accuracy | 0.945 |
| ECE (temperature 1.23) | 0.006 |
| coverage at 5% error | 0.99 |
effect accuracy |
0.828 |
| public decision-v7 regression (417 records) | 0.867 (Kev-4B: 0.862) |
Training
- Commands: 26,301 records.
- 23,674 are rippy q3 states built from tldr-pages examples (CC BY 4.0), three fills per example. Only commands rippy itself would send for review are kept.
- 2,627 are hard cases generated by Claude Sonnet: steering and lookalikes,
exfiltration, secret reads, download-and-run, destructive commands with
$VAR, benign commands that sound scary, and the read/write boundary. - The split is by program. Programs in rippy's gold sample are excluded everywhere.
- Labels: soft targets from weighted votes.
- Two independent Claude Sonnet passes label every record; they agree on all seven answers for 80% of records.
- Gemma 4 31B adds a third vote on dev and test.
- A blind Claude Opus adjudication, counted double, covers the 2,206 records whose teachers split on approval.
- Hosted Jev was only an evaluation reference, never a teacher.
- Recipe: Kev's delta fine-tune from
jaredpalmer/kev-4b.- 1 epoch, learning rate 2e-5, LoRA rank 16.
- 2,000 replay records from Kev's public recipe.
- About 2.3 hours on one H100.
- Temperature fitted on 3,113 calibration records.
Limitations
- Question set. Trained on rippy's
q3, where facts read e.g. "not found on rippy's PATH; may still exist when the command runs". Re-check a later question-set change with rippy'sscripts/jev-evalbefore relying on it. - Thresholds. Thresholds tuned for Jev do not carry over. Use the ones above, or
fit your own with rippy-kev's
eval/fit_thresholds.py. - Labels are model-made. Labels come from teacher models with a hand-labelled gold check, not from human review of every record. "Severe" and "safe" in the results are teacher verdicts.
- Remaining misses at the balanced setting are mostly config or settings
commands of unfamiliar tools (
blackfire config,mods --settings) and editors or viewers pointed at files outside the project (kak /etc/…). It also approved commands that print stored credentials (xauth list,mc alias list). The strict setting leaves 2 severe approvals per ~945 severe never-seen commands.
Licence and attribution
Apache-2.0. It builds on:
- Kev (Apache-2.0)
- Qwen3.5-4B-Base (Apache-2.0)
- commands derived from tldr-pages (CC BY 4.0)
- Downloads last month
- 42
8-bit
Model tree for Risethagain/rippy-kev-4b
Base model
Qwen/Qwen3.5-4B-Base