Instructions to use KucLab/kuclab-hertz-0.8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use KucLab/kuclab-hertz-0.8 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf KucLab/kuclab-hertz-0.8:Q4_K_M # Run inference directly in the terminal: llama cli -hf KucLab/kuclab-hertz-0.8:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf KucLab/kuclab-hertz-0.8:Q4_K_M # Run inference directly in the terminal: llama cli -hf KucLab/kuclab-hertz-0.8:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf KucLab/kuclab-hertz-0.8:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf KucLab/kuclab-hertz-0.8:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf KucLab/kuclab-hertz-0.8:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf KucLab/kuclab-hertz-0.8:Q4_K_M
Use Docker
docker model run hf.co/KucLab/kuclab-hertz-0.8:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use KucLab/kuclab-hertz-0.8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "KucLab/kuclab-hertz-0.8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "KucLab/kuclab-hertz-0.8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/KucLab/kuclab-hertz-0.8:Q4_K_M
- Ollama
How to use KucLab/kuclab-hertz-0.8 with Ollama:
ollama run hf.co/KucLab/kuclab-hertz-0.8:Q4_K_M
- Unsloth Desktop
- Pi
How to use KucLab/kuclab-hertz-0.8 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf KucLab/kuclab-hertz-0.8:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "KucLab/kuclab-hertz-0.8:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use KucLab/kuclab-hertz-0.8 with Docker Model Runner:
docker model run hf.co/KucLab/kuclab-hertz-0.8:Q4_K_M
- Lemonade
How to use KucLab/kuclab-hertz-0.8 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull KucLab/kuclab-hertz-0.8:Q4_K_M
Run and chat with the model
lemonade run user.kuclab-hertz-0.8-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use KucLab/kuclab-hertz-0.8 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf KucLab/kuclab-hertz-0.8:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default KucLab/kuclab-hertz-0.8:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use KucLab/kuclab-hertz-0.8 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf KucLab/kuclab-hertz-0.8:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "KucLab/kuclab-hertz-0.8:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
KucLab Hertz 0.8
Measured on this project's harness — same 240 MMLU-Pro STEM and 206 CZ terminology questions for every model.
A Czech/English STEM + programming assistant built by KucLab on top of Qwen/Qwen3.5-9B. Second-generation distillation: API teachers (DeepSeek-flash for STEM, Gemini 3.8 Flash for prose/Czech) with enforced answer-first discipline, trained with the proven Hertz recipe. It beats Hertz 0.7F on MMLU-Pro STEM.
What this is
Hertz 0.8 is a LoRA fine-tune (r=16, merged into the base weights) of Qwen3.5-9B:
- Base: Qwen/Qwen3.5-9B (~9B params, Apache 2.0)
- Method: QLoRA, r=16 / alpha=32, merged to bf16 then quantized
- Teachers: DeepSeek-flash via API (math, physics, chemistry, programming — reasoning, temperature 0.2–0.3) and Gemini 3.8 Flash via API (biology, design, Czech language, concise everyday answers — thinking disabled for discipline)
- Training data: ~5080 rows total — API-distilled STEM/prose/concise rows (answer-first, canonical
Answer: (X)final line on multiple-choice), 309 CS↔EN scientific-terminology rows with definitions (train split only), 28 answer-first formatting examples, identity rows (Hertz 0.8). Validated: 0 malformed, 0 duplicates. - Context: 32768 tokens in the Ollama Modelfile (
num_ctx). - Format available: GGUF (q4_k_m, ~5.3GB) for
llama.cpp/Ollama.
Quickstart (Ollama)
Important: ollama pull hf.co/... alone does NOT apply this model's system prompt. Use ollama create with the Modelfile below instead:
curl -O https://huggingface.co/KucLab/kuclab-hertz-0.8/resolve/main/Modelfile
ollama create kuclab-hertz-0.8 -f Modelfile
ollama run kuclab-hertz-0.8
Development story
Hertz 0.7F (83.3% STEM) was distilled from a local 30B teacher. For 0.8 we moved teachers to API frontier models and fixed two data diseases found by audit: (1) ~370 truncated answers cut by token limits were removed — they teach broken patterns; (2) parametrized templates with absurd ranges (93kW kettles, 862 bpm heart rates) were constrained to realistic values. A brevity domain (400 short Q&A) teaches the model that "ahoj" deserves one sentence, not an essay — same spirit as ThinkingCap-style token efficiency. Qwen3.5 still reasons by architecture; append "think": false to API requests for instant answers.
Benchmarks
Same prompts, same grading code, same Ollama Q4_K_M quantization, identical methodology throughout.
MMLU-Pro STEM (240 held-out questions, this project's own curated subset)
| Hertz 0.6 (12B) | Hertz 0.7F (9B) | Hertz 0.8 (9B) | |
|---|---|---|---|
| Biology | 91.7% | 83.3% | 86.7% |
| Chemistry | 61.7% | 78.3% | 83.3% |
| Math | 90.0% | 96.7% | 95.0% |
| Physics | 73.3% | 75.0% | 90.0% |
| Total | 79.2% | 83.3% | 88.8% |
Hertz 0.8 beats 0.7F by +5.5pp, with physics jumping +15pp. Only 10/240 answers needed the fallback re-ask (vs 22 in 0.7F).
Czech terminology benchmark (206 held-out CS↔EN scientific terms)
| Hertz 0.6 | Hertz 0.7F | Hertz 0.8 | |
|---|---|---|---|
| CS→EN | 82.5% | 88.3% | 87.4% |
| EN→CS | 65.0% | 82.5% | 81.6% |
| Total | 73.8% | 85.4% | 84.5% |
Essentially tied with 0.7F (−0.9pp) — the API-teacher Czech prose held the gains.
Honest status
- ✅ MMLU-Pro STEM: 88.8%, beats Hertz 0.7F (83.3%)
- ✅ Physics 90.0%, chemistry 83.3% — the API STEM distillation worked
- ✅ Correctly identifies as KucLab Hertz 0.8 (kuclab.org), no founder named
- ⚠️ Czech terminology (84.5%) essentially tied with 0.7F, not above it
- ⚠️ 91% STEM target not yet reached (−2.2pp) — next iteration
- ⏳ No tool-calling fine-tuning (tool-use rows were cut with the budget)
License
Apache 2.0, inherited from Qwen/Qwen3.5-9B.
Credits
- Base model: Qwen/Qwen3.5-9B (Qwen, Apache 2.0)
- Teachers: DeepSeek-flash (STEM) and Gemini 3.8 Flash (prose/Czech) via API
- Fine-tuning, dataset construction, and packaging: KucLab
- Downloads last month
- 13
4-bit
