Instructions to use bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M
- Ollama
How to use bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF with Ollama:
ollama run hf.co/bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF with Docker Model Runner:
docker model run hf.co/bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M
- Lemonade
How to use bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.ThinkingCap-Qwen3.6-27B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
base_model: bottlecapai/ThinkingCap-Qwen3.6-27B
base_model_relation: quantized
library_name: gguf
tags:
- qwen3_6
- gguf
- llama.cpp
- token-efficient
- efficient-thinking
pipeline_tag: image-text-to-text
extra_gated_heading: Request access to the BottleCap AI model
extra_gated_prompt: >-
Please fill out the form below to request access. Name, Company Name, and
Company Email are optional — just enter NA if you'd prefer not to share. Your
email may be used to send you information about BottleCap AI model updates and
early access before public release.
extra_gated_fields:
Name: text
Company Name: text
Company Email: text
extra_gated_button_content: Agree and request access
bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF
GGUF / llama.cpp quantizations of bottlecapai/ThinkingCap-Qwen3.6-27B — capability of Qwen3.6-27B with 50% less thinking tokens on average, achieved by finetuning Qwen3.6-27B (Qwen Team, 2026) while preserving the original answer quality and style.
➡️ Full model description, evaluation results (multi-seed, statistically tested), recommended sampling params, and citation: see the main model card at bottlecapai/ThinkingCap-Qwen3.6-27B.
About GGUF and quantization
GGUF is a single-file model format for running LLMs locally with llama.cpp and compatible runtimes (Ollama, LM Studio, …). The quantized variants below store weights at reduced precision — e.g. ≈4.7 bits per weight for Q4_K_M instead of the 16-bit f16 source — cutting download size and memory severalfold at a small, measured quality cost.
Files
| File | Quant | Size |
|---|---|---|
ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf |
Q4_K_M | 15.7 GB |
ThinkingCap-Qwen3.6-27B-Q6_K.gguf |
Q6_K | 20.9 GB |
ThinkingCap-Qwen3.6-27B-Q8_0.gguf |
Q8_0 | 27.1 GB |
ThinkingCap-Qwen3.6-27B-f16.gguf |
f16 | 50.9 GB |
mmproj-ThinkingCap-Qwen3.6-27B-f16.gguf |
mmproj (vision) | 0.9 GB |
f16 is the unquantized source; Q8_0 is near-lossless; Q6_K is a slightly smaller near-lossless option; Q4_K_M is the recommended size/quality balance for most local setups.
Usage (llama.cpp)
# pull a specific quant straight from the Hub and chat
llama-cli -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M -p "Hi"
# or download one file and run it
huggingface-cli download bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf --local-dir .
llama-cli -m ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf -p "Hi"
Speculative decoding (MTP)
llama.cpp can run MTP (multi-token-prediction) self-speculative decoding on these GGUFs for a decode speed-up — no separate draft model needed. Add --spec-type draft-mtp when serving:
llama-server -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M --spec-type draft-mtp
Set the draft length with --spec-draft-n-max (e.g. 4). Requires a recent llama.cpp build with MTP support.
Vision (image input)
ThinkingCap is a vision-language model. Image input needs the multimodal projector
mmproj-ThinkingCap-Qwen3.6-27B-f16.gguf (in this repo) loaded alongside a text GGUF — the
single f16 mmproj pairs with any of the quants above.
- LM Studio / Jan / Ollama, …: download the
mmproj-*.gguffrom this repo; LM Studio auto-detects it and enables the image (🖼️) button. - llama.cpp CLI:
huggingface-cli download bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF \
ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf mmproj-ThinkingCap-Qwen3.6-27B-f16.gguf --local-dir .
llama-mtmd-cli -m ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf \
--mmproj mmproj-ThinkingCap-Qwen3.6-27B-f16.gguf --image photo.jpg -p "Describe this image."
- llama-server: add
--mmproj mmproj-ThinkingCap-Qwen3.6-27B-f16.ggufto expose an OpenAI-compatible vision endpoint.
Expected performance
Measured on our serving harness (llama.cpp, CUDA llama-server) over MMLU-Pro (reasoning) and RealWorldQA (vision), N=200/dataset × 3 seeds, sampled decoding (temp 1.0 / top_p 0.95 / top_k 20). acc is the mean ± 95% CI across seeds; task s / tok/s are measured per request during the eval at batch size 16 (a throughput setting — a single local user decoding one request at a time will see meaningfully higher tok/s, but the per-row comparison is unaffected). For the full multi-benchmark evaluation see the main model card.
Accuracy is statistically identical across f16, the quantized variants, and the base model — every 95% CI overlaps (MMLU-Pro ≈0.89–0.91, RealWorldQA ≈0.78–0.82). Quantization is near-lossless, and the finetune matches base accuracy. What these GGUFs buy you is brevity → speed: the finetune reasons ~2× shorter than base (≈1000 vs ≈2100 MMLU-Pro tokens; ≈330 vs ≈700 RealWorldQA), so at equal accuracy it finishes each task far faster. MTP self-speculative decoding (--spec-type draft-mtp, n=4) accepts ≈3.3–3.7 tokens per verify step on top. The bolded row — Q4_K_M + MTP — is the fastest per task and the recommended size/quality balance; Q8_0 and Q6_K are near-lossless alternatives at larger sizes. For reference we also list unsloth's Dynamic GGUFs of the base model (UD-*): same llama.cpp path and same accuracy, but base-model quants reason ≈2× longer (none of the finetune's token savings), so they finish only ~1.2–2.2× faster than base-std vs our ~2.6–3.9×.
median tokens = median completion length; task s = measured end-to-end wall-clock per request (bs=16); tok/s = per-request decode rate under that load; speedup = base·bf16·std task s ÷ row task s — same llama.cpp path for every row, so the within-table comparison is apples-to-apples.
MMLU-Pro (reasoning)
| config | acc (mean ± 95% CI) | median tokens | tok/s | task s | speedup | accept_len (n=4) |
|---|---|---|---|---|---|---|
| Qwen3.6-27B base (bf16 GGUF) · standard | 0.908 ± 0.007 | 2098 | 16.0 | 138.1 | 1.00× | — |
| f16 · standard | 0.908 ± 0.038 | 1027 | 16.0 | 67.0 | 2.06× | — |
| f16 · MTP | 0.903 ± 0.007 | 1045 | 25.9 | 43.3 | 3.19× | 3.69 |
| Q8_0 · standard | 0.897 ± 0.044 | 990 | 21.0 | 49.6 | 2.78× | — |
| Q8_0 · MTP | 0.900 ± 0.012 | 1043 | 31.3 | 36.7 | 3.77× | 3.69 |
| Q6_K · standard | 0.908 ± 0.007 | 1011 | 22.1 | 49.2 | 2.81× | — |
| Q6_K · MTP | 0.910 ± 0.012 | 992 | 30.4 | 36.6 | 3.77× | 3.69 |
| Q4_K_M · standard | 0.903 ± 0.019 | 996 | 25.5 | 42.3 | 3.27× | — |
| Q4_K_M · MTP | 0.910 ± 0.012 | 1054 | 32.8 | 35.5 | 3.89× | 3.65 |
| unsloth UD-Q8_K_XL (base) · standard | 0.892 ± 0.014 | 2139 | 19.6 | 112.9 | 1.22× | — |
| unsloth UD-Q8_K_XL (base) · MTP | 0.905 ± 0.000 | 2133 | 31.2 | 72.3 | 1.91× | 3.66 |
| unsloth UD-Q4_K_XL (base) · standard | 0.900 ± 0.025 | 2172 | 26.0 | 86.3 | 1.60× | — |
| unsloth UD-Q4_K_XL (base) · MTP | 0.903 ± 0.014 | 2094 | 35.4 | 63.4 | 2.18× | 3.61 |
RealWorldQA (vision)
| config | acc (mean ± 95% CI) | median tokens | tok/s | task s | speedup | accept_len (n=4) |
|---|---|---|---|---|---|---|
| Qwen3.6-27B base (bf16 GGUF) · standard | 0.797 ± 0.019 | 700 | 15.1 | 50.9 | 1.00× | — |
| f16 · standard | 0.815 ± 0.045 | 340 | 13.4 | 29.3 | 1.74× | — |
| f16 · MTP | 0.818 ± 0.031 | 331 | 16.1 | 26.6 | 1.92× | 3.29 |
| Q8_0 · standard | 0.815 ± 0.054 | 318 | 16.5 | 23.8 | 2.14× | — |
| Q8_0 · MTP | 0.797 ± 0.019 | 335 | 21.2 | 21.0 | 2.42× | 3.32 |
| Q6_K · standard | 0.813 ± 0.007 | 359 | 17.0 | 25.0 | 2.04× | — |
| Q6_K · MTP | 0.817 ± 0.031 | 355 | 20.4 | 23.3 | 2.19× | 3.32 |
| Q4_K_M · standard | 0.808 ± 0.014 | 334 | 19.2 | 21.7 | 2.35× | — |
| Q4_K_M · MTP | 0.810 ± 0.045 | 315 | 21.3 | 19.8 | 2.57× | 3.26 |
| unsloth UD-Q8_K_XL (base) · standard | 0.795 ± 0.000 | 662 | 18.2 | 45.0 | 1.13× | — |
| unsloth UD-Q8_K_XL (base) · MTP | 0.782 ± 0.073 | 672 | 25.6 | 35.5 | 1.44× | 3.43 |
| unsloth UD-Q4_K_XL (base) · standard | 0.792 ± 0.031 | 727 | 23.5 | 40.2 | 1.27× | — |
| unsloth UD-Q4_K_XL (base) · MTP | 0.785 ± 0.081 | 760 | 29.1 | 35.3 | 1.44× | 3.38 |
