Instructions to use hacnho/gguf-final-logit-softcapping-backdoor-poc with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use hacnho/gguf-final-logit-softcapping-backdoor-poc with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf hacnho/gguf-final-logit-softcapping-backdoor-poc:Q2_K # Run inference directly in the terminal: llama cli -hf hacnho/gguf-final-logit-softcapping-backdoor-poc:Q2_K
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf hacnho/gguf-final-logit-softcapping-backdoor-poc:Q2_K # Run inference directly in the terminal: llama cli -hf hacnho/gguf-final-logit-softcapping-backdoor-poc:Q2_K
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf hacnho/gguf-final-logit-softcapping-backdoor-poc:Q2_K # Run inference directly in the terminal: ./llama-cli -hf hacnho/gguf-final-logit-softcapping-backdoor-poc:Q2_K
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf hacnho/gguf-final-logit-softcapping-backdoor-poc:Q2_K # Run inference directly in the terminal: ./build/bin/llama-cli -hf hacnho/gguf-final-logit-softcapping-backdoor-poc:Q2_K
Use Docker
docker model run hf.co/hacnho/gguf-final-logit-softcapping-backdoor-poc:Q2_K
- LM Studio
- Jan
- Ollama
How to use hacnho/gguf-final-logit-softcapping-backdoor-poc with Ollama:
ollama run hf.co/hacnho/gguf-final-logit-softcapping-backdoor-poc:Q2_K
- Unsloth Studio
How to use hacnho/gguf-final-logit-softcapping-backdoor-poc with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for hacnho/gguf-final-logit-softcapping-backdoor-poc to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for hacnho/gguf-final-logit-softcapping-backdoor-poc to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for hacnho/gguf-final-logit-softcapping-backdoor-poc to start chatting
- Atomic Chat new
- Docker Model Runner
How to use hacnho/gguf-final-logit-softcapping-backdoor-poc with Docker Model Runner:
docker model run hf.co/hacnho/gguf-final-logit-softcapping-backdoor-poc:Q2_K
- Lemonade
How to use hacnho/gguf-final-logit-softcapping-backdoor-poc with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull hacnho/gguf-final-logit-softcapping-backdoor-poc:Q2_K
Run and chat with the model
lemonade run user.gguf-final-logit-softcapping-backdoor-poc-Q2_K
List all available models
lemonade list
GGUF final_logit_softcapping output-manipulation PoC
This repository contains a benign security research proof of concept for a Huntr MFV report.
The malicious GGUF is the upstream Gemma4 tiny GGUF model with one non-tokenizer architecture metadata value changed:
gemma4.final_logit_softcapping: 30.0 -> 0.01
This value is loaded by llama.cpp into Gemma4 hyperparameters and applied to final logits during inference. With greedy sampling, the control model repeats prompt tokens, while the malicious model silently collapses generated tokens to <pad>.
Files:
gemma-4-1B-0.8B-tiny.Q2_K.final-logit-softcap-0.01.ggufreproduce.pyrequirements.txtbuild_poc.py
Control model:
Reproduction:
python3 -m venv /tmp/gguf-softcap-poc
. /tmp/gguf-softcap-poc/bin/activate
pip install -r requirements.txt
curl -L -o control.gguf \
https://huggingface.co/mradermacher/gemma-4-1B-0.8B-tiny-GGUF/resolve/main/gemma-4-1B-0.8B-tiny.Q2_K.gguf
curl -L -o malicious.gguf \
https://huggingface.co/hacnho/gguf-final-logit-softcapping-backdoor-poc/resolve/main/gemma-4-1B-0.8B-tiny.Q2_K.final-logit-softcap-0.01.gguf
LLAMA_SIMPLE=/path/to/llama-simple python reproduce.py control.gguf malicious.gguf
modelscan -p malicious.gguf --show-skipped
Expected result:
- control
Hello:<bos>HelloHelloHello... - malicious
Hello:<bos>Hello<pad><pad><pad>... modelscan==0.8.8:No issues foundand the GGUF is skipped
- Downloads last month
- 20
2-bit