Instructions to use bowmanslayer/VeriLoop-E2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use bowmanslayer/VeriLoop-E2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf bowmanslayer/VeriLoop-E2-GGUF:F16 # Run inference directly in the terminal: llama cli -hf bowmanslayer/VeriLoop-E2-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf bowmanslayer/VeriLoop-E2-GGUF:F16 # Run inference directly in the terminal: llama cli -hf bowmanslayer/VeriLoop-E2-GGUF:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf bowmanslayer/VeriLoop-E2-GGUF:F16 # Run inference directly in the terminal: ./llama-cli -hf bowmanslayer/VeriLoop-E2-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf bowmanslayer/VeriLoop-E2-GGUF:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf bowmanslayer/VeriLoop-E2-GGUF:F16
Use Docker
docker model run hf.co/bowmanslayer/VeriLoop-E2-GGUF:F16
- LM Studio
- Jan
- vLLM
How to use bowmanslayer/VeriLoop-E2-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bowmanslayer/VeriLoop-E2-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bowmanslayer/VeriLoop-E2-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/bowmanslayer/VeriLoop-E2-GGUF:F16
- Ollama
How to use bowmanslayer/VeriLoop-E2-GGUF with Ollama:
ollama run hf.co/bowmanslayer/VeriLoop-E2-GGUF:F16
- Unsloth Desktop
- Pi
How to use bowmanslayer/VeriLoop-E2-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bowmanslayer/VeriLoop-E2-GGUF:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "bowmanslayer/VeriLoop-E2-GGUF:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use bowmanslayer/VeriLoop-E2-GGUF with Docker Model Runner:
docker model run hf.co/bowmanslayer/VeriLoop-E2-GGUF:F16
- Lemonade
How to use bowmanslayer/VeriLoop-E2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull bowmanslayer/VeriLoop-E2-GGUF:F16
Run and chat with the model
lemonade run user.VeriLoop-E2-GGUF-F16
List all available models
lemonade list
- Hermes Agent
How to use bowmanslayer/VeriLoop-E2-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bowmanslayer/VeriLoop-E2-GGUF:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default bowmanslayer/VeriLoop-E2-GGUF:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use bowmanslayer/VeriLoop-E2-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bowmanslayer/VeriLoop-E2-GGUF:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "bowmanslayer/VeriLoop-E2-GGUF:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
VeriLoop-E2-GGUF
Community quantization of tsinghua-sigs-robot-lab/VeriLoop-E2, fixed revision 9379d199adcc44acdf836558037163878e3a37ae. This is not an official Tsinghua release. Model architecture is Qwen3_5ForConditionalGeneration; the upstream names the base Qwen3.8-27B.
What is included
Q4_K_M main model and separate F16 vision projector. No importance matrix was used. Full-precision conversion intermediates are excluded from this release.
No abliteration was applied to this version.
Local paired evaluation
Same frozen 1,514 questions, seed 20260816, thinking enabled with the original template and publisher raw-content protocol, greedy decoding, max output 32,768 tokens. These are our legacy sampled/custom scorers, not official leaderboard results. IFEval uses a partial custom rule implementation. See EVALUATION.md and evaluation.json for details and validity flags.
| Task | n | Original gguf | Ablated gguf | Delta pp |
|---|---|---|---|---|
| mmlu | 150 | 68.00% | 68.00% | +0.00 |
| cmmlu | 150 | 61.33% | 64.67% | +3.33 |
| mmlu_pro | 150 | 64.00% | 68.67% | +4.67 |
| ceval | 150 | 60.00% | 70.00% | +10.00 |
| arc | 150 | 90.00% | 91.33% | +1.33 |
| truthfulqa | 150 | 60.67% | 60.67% | +0.00 |
| gsm8k | 100 | 37.00% | 47.00% | +10.00 |
| math500 | 100 | 68.00% | 67.00% | -1.00 |
| bbh | 150 | 48.00% | 54.67% | +6.67 |
| humaneval | 164 | 90.85% | 90.24% | -0.61 |
| ifeval | 100 | 63.00% | 60.00% | -3.00 |
Overall original → ablated: 65.72% → 68.63%, difference +2.91 pp; approximate paired 95% CI [0.956, 4.856] pp. Empty responses: [4, 7]; length-limited responses: [4, 7]. Both stay in the denominator. Empty-rate validity gate: True.
On 100 separate refusal probes (thinking disabled, 2048-token budget), opening-30-word marker counts were 99/100 → 0/100. Marker absence is not proof of semantic compliance. Truncation/budget hits: [0, 7]; validity: [True, True]. The first 16 probes overlap candidate selection; see the 84-item held-out subset in evaluation.json when available.
Loading
llama-server -m original-Q4_K_M.gguf -ngl 99 -c 40960 --jinja --chat-template-file chat_template.jinja --reasoning-format none
Use a recent compatible llama.cpp build and the supplied template override: the GGUF metadata inherited an older identity block from the upstream tokenizer configuration, whereas the standalone upstream template is authoritative. All 1,514 evaluated prompts and token sequences were verified against the standalone source template. The separate mmproj-F16.gguf is included; this evaluation used text-only requests. Visual inference was not validated. Standard GGUF conversion excludes MTP and custom conditional-memory execution.
Scope and limitations
Use the repository-shipped template and publisher raw-content serving protocol. Do not add a Qwen3 reasoning parser: some completed direct answers omit a closing think tag and that parser classifies them entirely as hidden reasoning. Removing it recovered all 33 initial missing-answer cases; streaming/nonstreaming checks passed. Thinking remains enabled. For evaluation, final content is taken after </think> when present; normally stopped direct content is retained, while unclosed length-limited output is not scored as an answer.
In llama.cpp content-only mode, remove an echoed opening <think> generation-prefix tag before applying answer extraction. It is protocol framing, not answer text. The GGUF preflight was restarted after verifying this correction and the template override; it is excluded from the reported scores.
The author's private VeriLoop Harness is not included or reproduced. Their SWE/Terminal/DeepSWE scores are not scores for these weights. This evaluation does not establish long-context, tool-use, vision, agentic performance, or production suitability. Reduced refusal can affect inappropriate-content handling; applications need their own behavior controls. Full semantic refusal review and comprehensive capability equivalence have not been established.
Apache-2.0 is retained from the fixed upstream source; see LICENSE. Source training belongs to its original authors. Community changes are quantization. See release.json and SHA256SUMS for provenance and integrity.
Complete measurement records and historical comparison
See all six result columns, test standards and limitations, all 9,084 per-item metric records, question IDs and hashes, and 400 refusal-probe metric records. No full-precision baseline was measured. Historical results are not a strictly identical-runtime comparison.
- Downloads last month
- 65
4-bit