adasgaleus's picture
Update README.md
2ea7e9b verified
|
Raw
History Blame Contribute Delete
8.68 kB
metadata
base_model: bottlecapai/ThinkingCap-Qwen3.6-27B
base_model_relation: quantized
library_name: gguf
tags:
  - qwen3_6
  - gguf
  - llama.cpp
  - token-efficient
  - efficient-thinking
pipeline_tag: image-text-to-text
extra_gated_heading: Request access to the BottleCap AI model
extra_gated_prompt: >-
  Please fill out the form below to request access. Name, Company Name, and
  Company Email are optional — just enter NA if you'd prefer not to share. Your
  email may be used to send you information about BottleCap AI model updates and
  early access before public release.
extra_gated_fields:
  Name: text
  Company Name: text
  Company Email: text
extra_gated_button_content: Agree and request access

ThinkingCap — BottleCap AI

bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF

GGUF / llama.cpp quantizations of bottlecapai/ThinkingCap-Qwen3.6-27B — capability of Qwen3.6-27B with 50% less thinking tokens on average, achieved by finetuning Qwen3.6-27B (Qwen Team, 2026) while preserving the original answer quality and style.

➡️ Full model description, evaluation results (multi-seed, statistically tested), recommended sampling params, and citation: see the main model card at bottlecapai/ThinkingCap-Qwen3.6-27B.

About GGUF and quantization

GGUF is a single-file model format for running LLMs locally with llama.cpp and compatible runtimes (Ollama, LM Studio, …). The quantized variants below store weights at reduced precision — e.g. ≈4.7 bits per weight for Q4_K_M instead of the 16-bit f16 source — cutting download size and memory severalfold at a small, measured quality cost.

Files

File Quant Size
ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf Q4_K_M 15.7 GB
ThinkingCap-Qwen3.6-27B-Q6_K.gguf Q6_K 20.9 GB
ThinkingCap-Qwen3.6-27B-Q8_0.gguf Q8_0 27.1 GB
ThinkingCap-Qwen3.6-27B-f16.gguf f16 50.9 GB
mmproj-ThinkingCap-Qwen3.6-27B-f16.gguf mmproj (vision) 0.9 GB

f16 is the unquantized source; Q8_0 is near-lossless; Q6_K is a slightly smaller near-lossless option; Q4_K_M is the recommended size/quality balance for most local setups.

Usage (llama.cpp)

# pull a specific quant straight from the Hub and chat
llama-cli -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M -p "Hi"

# or download one file and run it
huggingface-cli download bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf --local-dir .
llama-cli -m ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf -p "Hi"

Speculative decoding (MTP)

llama.cpp can run MTP (multi-token-prediction) self-speculative decoding on these GGUFs for a decode speed-up — no separate draft model needed. Add --spec-type draft-mtp when serving:

llama-server -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M --spec-type draft-mtp

Set the draft length with --spec-draft-n-max (e.g. 4). Requires a recent llama.cpp build with MTP support.

Vision (image input)

ThinkingCap is a vision-language model. Image input needs the multimodal projector mmproj-ThinkingCap-Qwen3.6-27B-f16.gguf (in this repo) loaded alongside a text GGUF — the single f16 mmproj pairs with any of the quants above.

  • LM Studio / Jan / Ollama, …: download the mmproj-*.gguf from this repo; LM Studio auto-detects it and enables the image (🖼️) button.
  • llama.cpp CLI:
huggingface-cli download bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF \
  ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf mmproj-ThinkingCap-Qwen3.6-27B-f16.gguf --local-dir .
llama-mtmd-cli -m ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf \
  --mmproj mmproj-ThinkingCap-Qwen3.6-27B-f16.gguf --image photo.jpg -p "Describe this image."
  • llama-server: add --mmproj mmproj-ThinkingCap-Qwen3.6-27B-f16.gguf to expose an OpenAI-compatible vision endpoint.

Expected performance

Measured on our serving harness (llama.cpp, CUDA llama-server) over MMLU-Pro (reasoning) and RealWorldQA (vision), N=200/dataset × 3 seeds, sampled decoding (temp 1.0 / top_p 0.95 / top_k 20). acc is the mean ± 95% CI across seeds; task s / tok/s are measured per request during the eval at batch size 16 (a throughput setting — a single local user decoding one request at a time will see meaningfully higher tok/s, but the per-row comparison is unaffected). For the full multi-benchmark evaluation see the main model card.

Accuracy is statistically identical across f16, the quantized variants, and the base model — every 95% CI overlaps (MMLU-Pro ≈0.89–0.91, RealWorldQA ≈0.78–0.82). Quantization is near-lossless, and the finetune matches base accuracy. What these GGUFs buy you is brevity → speed: the finetune reasons ~2× shorter than base (≈1000 vs ≈2100 MMLU-Pro tokens; ≈330 vs ≈700 RealWorldQA), so at equal accuracy it finishes each task far faster. MTP self-speculative decoding (--spec-type draft-mtp, n=4) accepts ≈3.3–3.7 tokens per verify step on top. The bolded row — Q4_K_M + MTP — is the fastest per task and the recommended size/quality balance; Q8_0 and Q6_K are near-lossless alternatives at larger sizes. For reference we also list unsloth's Dynamic GGUFs of the base model (UD-*): same llama.cpp path and same accuracy, but base-model quants reason ≈2× longer (none of the finetune's token savings), so they finish only ~1.2–2.2× faster than base-std vs our ~2.6–3.9×.

median tokens = median completion length; task s = measured end-to-end wall-clock per request (bs=16); tok/s = per-request decode rate under that load; speedup = base·bf16·std task s ÷ row task s — same llama.cpp path for every row, so the within-table comparison is apples-to-apples.

MMLU-Pro (reasoning)

config acc (mean ± 95% CI) median tokens tok/s task s speedup accept_len (n=4)
Qwen3.6-27B base (bf16 GGUF) · standard 0.908 ± 0.007 2098 16.0 138.1 1.00×
f16 · standard 0.908 ± 0.038 1027 16.0 67.0 2.06×
f16 · MTP 0.903 ± 0.007 1045 25.9 43.3 3.19× 3.69
Q8_0 · standard 0.897 ± 0.044 990 21.0 49.6 2.78×
Q8_0 · MTP 0.900 ± 0.012 1043 31.3 36.7 3.77× 3.69
Q6_K · standard 0.908 ± 0.007 1011 22.1 49.2 2.81×
Q6_K · MTP 0.910 ± 0.012 992 30.4 36.6 3.77× 3.69
Q4_K_M · standard 0.903 ± 0.019 996 25.5 42.3 3.27×
Q4_K_M · MTP 0.910 ± 0.012 1054 32.8 35.5 3.89× 3.65
unsloth UD-Q8_K_XL (base) · standard 0.892 ± 0.014 2139 19.6 112.9 1.22×
unsloth UD-Q8_K_XL (base) · MTP 0.905 ± 0.000 2133 31.2 72.3 1.91× 3.66
unsloth UD-Q4_K_XL (base) · standard 0.900 ± 0.025 2172 26.0 86.3 1.60×
unsloth UD-Q4_K_XL (base) · MTP 0.903 ± 0.014 2094 35.4 63.4 2.18× 3.61

RealWorldQA (vision)

config acc (mean ± 95% CI) median tokens tok/s task s speedup accept_len (n=4)
Qwen3.6-27B base (bf16 GGUF) · standard 0.797 ± 0.019 700 15.1 50.9 1.00×
f16 · standard 0.815 ± 0.045 340 13.4 29.3 1.74×
f16 · MTP 0.818 ± 0.031 331 16.1 26.6 1.92× 3.29
Q8_0 · standard 0.815 ± 0.054 318 16.5 23.8 2.14×
Q8_0 · MTP 0.797 ± 0.019 335 21.2 21.0 2.42× 3.32
Q6_K · standard 0.813 ± 0.007 359 17.0 25.0 2.04×
Q6_K · MTP 0.817 ± 0.031 355 20.4 23.3 2.19× 3.32
Q4_K_M · standard 0.808 ± 0.014 334 19.2 21.7 2.35×
Q4_K_M · MTP 0.810 ± 0.045 315 21.3 19.8 2.57× 3.26
unsloth UD-Q8_K_XL (base) · standard 0.795 ± 0.000 662 18.2 45.0 1.13×
unsloth UD-Q8_K_XL (base) · MTP 0.782 ± 0.073 672 25.6 35.5 1.44× 3.43
unsloth UD-Q4_K_XL (base) · standard 0.792 ± 0.031 727 23.5 40.2 1.27×
unsloth UD-Q4_K_XL (base) · MTP 0.785 ± 0.081 760 29.1 35.3 1.44× 3.38