Instructions to use ressl/gemma-4-31B-it-uncensored-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ressl/gemma-4-31B-it-uncensored-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M
Use Docker
docker model run hf.co/ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use ressl/gemma-4-31B-it-uncensored-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ressl/gemma-4-31B-it-uncensored-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ressl/gemma-4-31B-it-uncensored-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M
- Ollama
How to use ressl/gemma-4-31B-it-uncensored-GGUF with Ollama:
ollama run hf.co/ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use ressl/gemma-4-31B-it-uncensored-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ressl/gemma-4-31B-it-uncensored-GGUF with Docker Model Runner:
docker model run hf.co/ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M
- Lemonade
How to use ressl/gemma-4-31B-it-uncensored-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.gemma-4-31B-it-uncensored-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use ressl/gemma-4-31B-it-uncensored-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ressl/gemma-4-31B-it-uncensored-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ressl/gemma-4-31B-it-uncensored-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
gemma-4-31B-it-uncensored (GGUF)
GGUF K-quant ladder of ressl/gemma-4-31B-it-uncensored for llama.cpp, Ollama, and LM Studio. Uncensored: 0-1/100 effective refusals at every quant level (the uncensoring survives quantization down to 2-bit).
⚠️ Genuinely uncensored, it will comply with requests a stock model refuses.
Intended use, the constructive side. A non-refusing assistant is genuinely useful for ethical hacking, security research, and penetration testing: red-teaming, analyzing malware and exploit code, writing detection/YARA rules, reviewing vulnerabilities, and studying attack techniques without the model bailing out mid-task. Use it lawfully and responsibly.
ℹ️ gemma-4 has a thinking mode. llama.cpp enables it by default, so the answer lands in
reasoning_contentandcontentcan look empty. For direct answers pass--reasoning-budget 0(llama-server) or disable thinking in your client.
Format set
| Repository | Format | Runs on |
|---|---|---|
| ressl/gemma-4-31B-it-uncensored | Transformers BF16, multimodal | transformers, vLLM, SGLang |
| ressl/gemma-4-31B-it-uncensored-NVFP4 | NVIDIA NVFP4, multimodal | vLLM, SGLang on Blackwell |
| ressl/gemma-4-31B-it-uncensored-GGUF | GGUF q8_0 to q2_k, text only | llama.cpp, Ollama, LM Studio |
| ressl/gemma-4-31B-it-uncensored-MLX-bf16 | MLX BF16, multimodal | mlx-vlm on Apple silicon |
| ressl/gemma-4-31B-it-uncensored-MLX-8bit | MLX 8-bit, multimodal | mlx-vlm on Apple silicon |
| ressl/gemma-4-31B-it-uncensored-MLX-6bit | MLX 6-bit, multimodal | mlx-vlm on Apple silicon |
| ressl/gemma-4-31B-it-uncensored-MLX-5bit | MLX 5-bit, multimodal | mlx-vlm on Apple silicon |
| ressl/gemma-4-31B-it-uncensored-MLX-4bit | MLX 4-bit, multimodal | mlx-vlm on Apple silicon |
Quants
Text-only (llama.cpp drops the vision tower, the BF16/NVFP4 repos keep multimodal).
| File | Size | Use when |
|---|---|---|
…-q8_0.gguf |
31 GB | maximum quality |
…-q6_k.gguf |
24 GB | near-lossless, smaller |
…-q5_k_m.gguf |
21 GB | high quality |
…-q4_k_m.gguf |
18 GB | recommended default, best size/quality |
…-q3_k_m.gguf |
15 GB | tight VRAM |
…-q2_k.gguf |
12 GB | smallest; still 0/100 uncensored, some quality loss |
For full precision use the BF16 repo (an
f16 GGUF exceeds Hugging Face's 50 GB per-file limit and is not hosted here).
All measured at 0/100 hard refusals except q5_k_m (1/100). GPU inference ~75 tok/s (q4, one RTX PRO 6000).
Run it with llama.cpp
llama-server -m gemma-4-31B-it-uncensored-biproj-q4_k_m.gguf \
-ngl 99 -c 8192 --reasoning-budget 0
Run it with Ollama
# Modelfile: FROM ./gemma-4-31B-it-uncensored-biproj-q4_k_m.gguf
ollama create gemma4-unc -f Modelfile && ollama run gemma4-unc
Needs a recent llama.cpp/Ollama build with Gemma4 GGUF support (build ≥ 2026-06).
Facts & figures
| Base | ressl/gemma-4-31B-it-uncensored → google/gemma-4-31B-it |
| Type | uncensored (abliterated) build, 0/686 effective refusals across 4 datasets |
| Converter | llama.cpp convert_hf_to_gguf.py + llama-quantize |
Cross-dataset validation
Generalization tested across 686 prompts from 4 independent datasets, 0 effective refusals everywhere:
| Dataset | Prompts | Effective refusals |
|---|---|---|
| JailbreakBench | 100 | 0/100 |
| tulu-harmbench | 320 | 0/320 |
| NousResearch/RefusalDataset | 166 | 0/166 |
| mlabonne/harmful_behaviors | 100 | 0/100 |
| Total | 686 | 0/686 (0.0%) |
A naive keyword detector flags 363/686 (52.9%), every one is a ***Disclaimer:**-prefixed
compliant answer, not a refusal. (Measured on the shared abliterated weights.)
❤️ Support
Producing and validating this complete format set (BF16 + NVFP4 + a full GGUF ladder, across vLLM, SGLang and llama.cpp on bleeding-edge Blackwell hardware) was a lot of work. If it's useful to you, I'd genuinely appreciate your support on Patreon 🙏, more at ressl.ch.
License & credits
Apache License 2.0, inherited from the base model by Google. See the official Gemma 4 license page. Uncensoring, format set and validation by Robert Ressl (Hugging Face · Website · LinkedIn · Patreon).
- Downloads last month
- 1,679
2-bit
3-bit
4-bit
5-bit
6-bit
8-bit
Model tree for ressl/gemma-4-31B-it-uncensored-GGUF
Base model
google/gemma-4-31B