Instructions to use haakona/Qwen-Unitopia-Style with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use haakona/Qwen-Unitopia-Style with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf haakona/Qwen-Unitopia-Style:Q4_K_M # Run inference directly in the terminal: llama cli -hf haakona/Qwen-Unitopia-Style:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf haakona/Qwen-Unitopia-Style:Q4_K_M # Run inference directly in the terminal: llama cli -hf haakona/Qwen-Unitopia-Style:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf haakona/Qwen-Unitopia-Style:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf haakona/Qwen-Unitopia-Style:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf haakona/Qwen-Unitopia-Style:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf haakona/Qwen-Unitopia-Style:Q4_K_M
Use Docker
docker model run hf.co/haakona/Qwen-Unitopia-Style:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use haakona/Qwen-Unitopia-Style with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "haakona/Qwen-Unitopia-Style" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "haakona/Qwen-Unitopia-Style", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/haakona/Qwen-Unitopia-Style:Q4_K_M
- Ollama
How to use haakona/Qwen-Unitopia-Style with Ollama:
ollama run hf.co/haakona/Qwen-Unitopia-Style:Q4_K_M
- Unsloth Desktop
- Pi
How to use haakona/Qwen-Unitopia-Style with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf haakona/Qwen-Unitopia-Style:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "haakona/Qwen-Unitopia-Style:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use haakona/Qwen-Unitopia-Style with Docker Model Runner:
docker model run hf.co/haakona/Qwen-Unitopia-Style:Q4_K_M
- Lemonade
How to use haakona/Qwen-Unitopia-Style with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull haakona/Qwen-Unitopia-Style:Q4_K_M
Run and chat with the model
lemonade run user.Qwen-Unitopia-Style-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use haakona/Qwen-Unitopia-Style with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf haakona/Qwen-Unitopia-Style:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default haakona/Qwen-Unitopia-Style:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use haakona/Qwen-Unitopia-Style with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf haakona/Qwen-Unitopia-Style:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "haakona/Qwen-Unitopia-Style:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen-Unitopia-Style
LoRA fine-tunes of Qwen models on the UNItopia LPC mudlib, a German
LPMud library (sources and documentation from ftp://unitopia.de). The goal is
a model that writes LPC in the UNItopia house style: file headers, inherit
lines, the room, item and monster APIs, and the German documentation format.
The current release is built on Qwen3.8-27B.
Files
Everything is in qwen3.8-27b/:
| File | Size | What |
|---|---|---|
qwen3.8-27b-unitopia-q8_0.gguf |
29.0 GB | The fine-tune. Q8_0, near-lossless, includes the multi-token-prediction block |
qwen3.8-27b-unitopia-q4_k_m.gguf |
16.8 GB | The fine-tune in plain Q4_K_M (no importance matrix), for smaller cards |
qwen3.8-27b-base-q8_0.gguf |
29.0 GB | The unmodified base model in the same Q8_0 conversion, for A/B comparison |
adapter-pretrain/, adapter-sft/ |
435 MB each | PEFT LoRA adapters of the two phases (rank 16, alpha 32, 109M params). adapter-sft is what the merged GGUFs contain; load it on Qwen/Qwen3.8-27B with PeftModel.from_pretrained |
metrics-*.jsonl, run-*.json |
Training curves and resolved arguments | |
SHA256SUMS |
Checksums of the three GGUFs |
All GGUFs are text-only conversions with llama.cpp (convert_hf_to_gguf.py);
the vision tower of the base checkpoint is not included. The Q8_0 files were
converted directly from bf16 safetensors, the Q4_K_M with llama-quantize
at default settings.
Also here: Qwen3.5-35B-A3B ("flash-lite")
qwen3.5-35b-a3b/ holds the same fine-tune applied to Qwen3.5-35B-A3B, a
mixture-of-experts model with 3B active parameters: qwen3.5-35b-a3b-unitopia-q8_0.gguf
(37.8 GB, Q8_0 with MTP, GGUF architecture qwen35moe), the two LoRA adapters
and the training metrics. It answers at roughly the speed of a 3B model with
the knowledge of a 35B one; final eval loss 1.14 against the 27B's 0.97, so
the 27B writes better code and the 35B-A3B answers faster.
Also here: Gemma 4 31B
gemma-4-31b/ holds the same fine-tune on Google's Gemma 4 31B
(google/gemma-4-31B-it, dense, Apache 2.0), trained on one H200 on
2026-10-02:
| File | Size | What |
|---|---|---|
gemma-4-31b-unitopia-q8_0.gguf |
32.6 GB | The fine-tune, Q8_0 |
gemma-4-31b-unitopia-q4_k_m.gguf |
18.7 GB | The fine-tune in plain Q4_K_M (no importance matrix) |
gemma-4-31b-unitopia-bf16.gguf |
61.4 GB | The fine-tune unquantized (bf16), the source for your own quantizations |
gemma-4-31b-base-q8_0.gguf |
32.6 GB | The unmodified base model in Q8_0, for A/B comparison |
adapter-pretrain/, adapter-sft/ |
490 MB each | PEFT LoRA adapters (rank 16, alpha 32, 122M params) for google/gemma-4-31B-it |
metrics-*.jsonl, run-*.json, SHA256SUMS |
Training curves, resolved arguments, checksums of the four GGUFs |
Same two phases and settings as the Qwen runs: phase 1 (causal LM on the
mudlib) 139 min, eval loss 6.22 → 1.03; phase 2 (instruction pairs) 38 min,
1.10 → 1.02. The high phase-1 start is the instruction-tuned Gemma's poor fit
to raw non-chat text (llama.cpp measures the same on the base model), and
phase 1 removes it within a hundred steps. Training used the chat template
with thinking off; Gemma 4 switches thinking on with <|think|> in the
system prompt (enable_thinking), and the model works either way. The
GGUFs are text-only (GGUF architecture gemma4) and carry Gemma's chat
template.
How it was trained
Two LoRA phases with PyTorch and PEFT on one RTX PRO 6000 (96 GB):
| Phase | Data | Optimizer steps | Time | Eval loss start → end |
|---|---|---|---|---|
| 1 causal LM | mudlib sources and docs, 4411 windows of ≤2048 tokens, 3.97M target tokens, 2 epochs | 1102 | 274 min | 1.51 → 0.97 |
| 2 SFT | 2815 instruction pairs derived from the sources (write this file, explain this help page, where is X), 2 epochs | 703 | 78 min | 1.02 → 0.97 |
LoRA on all attention, MLP and gated-delta-net projections, bf16 base, gradient checkpointing, AdamW lr 2e-4 with warmup and cosine decay. Eval loss is measured on a held-out 5% of the same data. For scale, the same recipe gave 1.38 on Qwen3-0.6B, 1.19 on Qwen3-4B-Instruct-2507 and 1.06 on Qwen3.5-9B.
Using it
Prompts in German work best, phrased like the training data:
Implementiere `/room/kirche/treppe5.c` für die UNItopia Mudlib (Die Treppe zum Kirchturm).
Dokumentation für die Hilfeseite `rm` (UNItopia Mudlib)?
Schreibe einen einfachen NPC für die UNItopia Mudlib: ein Bäcker, der Brot verkauft und auf 'hallo' antwortet.
Training used the chat template with thinking disabled (enable_thinking=False);
the model still works with thinking on. Tool calling from the base model is
intact. With llama.cpp:
llama-cli -m qwen3.8-27b-unitopia-q8_0.gguf -ngl 99 -cnv
What to expect
It writes idiomatic UNItopia LPC, reproduces the documentation format, and
composes new objects from the mudlib's conventions (a seller inherits
/i/money/verkaeufer, an NPC /i/monster/monster with monster::create()).
It is a junior builder, not a senior one: when it does not remember a
function it tends to invent a plausible name instead of saying so. Give it
file access to the mudlib and a system prompt that tells it where the
mudlib lives and to look functions up before using them, and review what it
writes before anything reaches players.
Licences
The base models are Qwen3.8-27B and Qwen3.5-35B-A3B under their own licences and Gemma 4 31B under Apache 2.0. The training data is the UNItopia mudlib and documentation, which are licensed for non-commercial use only; treat these fine-tunes the same way. Base weights are redistributed here only as a conversion for comparison.
- Downloads last month
- 364
4-bit
8-bit
16-bit
Model tree for haakona/Qwen-Unitopia-Style
Base model
Qwen/Qwen3.8-27B