VeriLoop-E2-GGUF

Community quantization of tsinghua-sigs-robot-lab/VeriLoop-E2, fixed revision 9379d199adcc44acdf836558037163878e3a37ae. This is not an official Tsinghua release. Model architecture is Qwen3_5ForConditionalGeneration; the upstream names the base Qwen3.8-27B.

What is included

Q4_K_M main model and separate F16 vision projector. No importance matrix was used. Full-precision conversion intermediates are excluded from this release.

No abliteration was applied to this version.

Local paired evaluation

Same frozen 1,514 questions, seed 20260816, thinking enabled with the original template and publisher raw-content protocol, greedy decoding, max output 32,768 tokens. These are our legacy sampled/custom scorers, not official leaderboard results. IFEval uses a partial custom rule implementation. See EVALUATION.md and evaluation.json for details and validity flags.

Task n Original gguf Ablated gguf Delta pp
mmlu 150 68.00% 68.00% +0.00
cmmlu 150 61.33% 64.67% +3.33
mmlu_pro 150 64.00% 68.67% +4.67
ceval 150 60.00% 70.00% +10.00
arc 150 90.00% 91.33% +1.33
truthfulqa 150 60.67% 60.67% +0.00
gsm8k 100 37.00% 47.00% +10.00
math500 100 68.00% 67.00% -1.00
bbh 150 48.00% 54.67% +6.67
humaneval 164 90.85% 90.24% -0.61
ifeval 100 63.00% 60.00% -3.00

Overall original → ablated: 65.72% → 68.63%, difference +2.91 pp; approximate paired 95% CI [0.956, 4.856] pp. Empty responses: [4, 7]; length-limited responses: [4, 7]. Both stay in the denominator. Empty-rate validity gate: True.

On 100 separate refusal probes (thinking disabled, 2048-token budget), opening-30-word marker counts were 99/100 → 0/100. Marker absence is not proof of semantic compliance. Truncation/budget hits: [0, 7]; validity: [True, True]. The first 16 probes overlap candidate selection; see the 84-item held-out subset in evaluation.json when available.

Loading

llama-server -m original-Q4_K_M.gguf -ngl 99 -c 40960 --jinja --chat-template-file chat_template.jinja --reasoning-format none

Use a recent compatible llama.cpp build and the supplied template override: the GGUF metadata inherited an older identity block from the upstream tokenizer configuration, whereas the standalone upstream template is authoritative. All 1,514 evaluated prompts and token sequences were verified against the standalone source template. The separate mmproj-F16.gguf is included; this evaluation used text-only requests. Visual inference was not validated. Standard GGUF conversion excludes MTP and custom conditional-memory execution.

Scope and limitations

Use the repository-shipped template and publisher raw-content serving protocol. Do not add a Qwen3 reasoning parser: some completed direct answers omit a closing think tag and that parser classifies them entirely as hidden reasoning. Removing it recovered all 33 initial missing-answer cases; streaming/nonstreaming checks passed. Thinking remains enabled. For evaluation, final content is taken after </think> when present; normally stopped direct content is retained, while unclosed length-limited output is not scored as an answer.

In llama.cpp content-only mode, remove an echoed opening <think> generation-prefix tag before applying answer extraction. It is protocol framing, not answer text. The GGUF preflight was restarted after verifying this correction and the template override; it is excluded from the reported scores.

The author's private VeriLoop Harness is not included or reproduced. Their SWE/Terminal/DeepSWE scores are not scores for these weights. This evaluation does not establish long-context, tool-use, vision, agentic performance, or production suitability. Reduced refusal can affect inappropriate-content handling; applications need their own behavior controls. Full semantic refusal review and comprehensive capability equivalence have not been established.

Apache-2.0 is retained from the fixed upstream source; see LICENSE. Source training belongs to its original authors. Community changes are quantization. See release.json and SHA256SUMS for provenance and integrity.

Complete measurement records and historical comparison

See all six result columns, test standards and limitations, all 9,084 per-item metric records, question IDs and hashes, and 400 refusal-probe metric records. No full-precision baseline was measured. Historical results are not a strictly identical-runtime comparison.

Downloads last month
65
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bowmanslayer/VeriLoop-E2-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(12)
this model