Instructions to use GestaltLabs/Ornstein3.8-27B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use GestaltLabs/Ornstein3.8-27B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use GestaltLabs/Ornstein3.8-27B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "GestaltLabs/Ornstein3.8-27B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GestaltLabs/Ornstein3.8-27B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M
- Ollama
How to use GestaltLabs/Ornstein3.8-27B-GGUF with Ollama:
ollama run hf.co/GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use GestaltLabs/Ornstein3.8-27B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use GestaltLabs/Ornstein3.8-27B-GGUF with Docker Model Runner:
docker model run hf.co/GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M
- Lemonade
How to use GestaltLabs/Ornstein3.8-27B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Ornstein3.8-27B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use GestaltLabs/Ornstein3.8-27B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use GestaltLabs/Ornstein3.8-27B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "GestaltLabs/Ornstein3.8-27B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Ornstein3.8-27B-GGUF
GGUF quantizations of GestaltLabs/Ornstein3.8-27B: a Qwen 3.8 27B dense vision-language fine-tune with interleaved linear and full attention (Gated DeltaNet). The LoRA was trained on Fireworks AI and merged into the Qwen3.8-27B language stack.
Language GGUFs are qwen35 text-path only (866 tensors). For llama.cpp vision, pass --mmproj mmproj-Ornstein3.8-27B-BF16.gguf. For Transformers or vLLM multimodal inference, use the base safetensors.
Status
This checkpoint injects Ornstein thinking into Qwen3.8-27B. It is an early merge, not a finished quality release. Planned quality work uses RL environments and energy-based fine-tuning.
Evaluation
Qwen3.8-27B achieves an estimated 97.0% accuracy on the full GSM8K benchmark when running in standard unquantized precision (BF16/FP8).
| Benchmark | Qwen3.8-27B (reported) | Ornstein3.8-27B (this run) |
|---|---|---|
| GSM8K | — | 96.51 (1273/1319) |
Single greedy BF16 run on a Fireworks dedicated H100 (temperature=0, top_k=40, max_tokens=4000), answers from message.content first. Qwen does not report GSM8K on the Qwen3.8-27B card. This score does not apply to the GGUF files in this repo.
Support this work
I'm a PhD student in visual neuroscience at the University of Toronto. Training and release compute is self-funded (rented H100s and a local DGX Spark). If these artifacts are useful, Ko-fi helps keep the experiments running.
Model details
| Architecture | Qwen3_5ForConditionalGeneration |
| Parameters | ~27B dense |
| Context | 262,144 tokens |
| Hidden size / layers | 5120 / 64 |
| Attention | 24 heads, 4 KV heads, head_dim 256 |
| Training | rank-32 LoRA on Fireworks AI |
| Language GGUF tensors | 866 (token_embd, blk, output) |
| mmproj tensors | 334 (v.* vision tower, mm.* projector) |
Quant index
Pick a quant that fits in RAM/VRAM with room for context. For a dense 27B, Q4_K_M or higher is the usual 24 GB default; use Q6_K when you have headroom.
| File | Size | Bits | Notes |
|---|---|---|---|
Ornstein3.8-27B-Q8_0.gguf |
29.0 GB | 8 | Near-lossless reference |
Ornstein3.8-27B-Q6_K.gguf |
22.4 GB | 6.5 | Strong default on 32 GB+ |
Ornstein3.8-27B-Q4_K_M.gguf |
16.8 GB | 4.5 | Common 24 GB default |
mmproj-Ornstein3.8-27B-BF16.gguf |
931 MB | BF16 | Vision tower + projector; required for image/video in llama.cpp |
BF16 language GGUF is not shipped here. Full precision is the safetensors in the base repo.
Usage
Requires llama.cpp with qwen35 / Qwen3.8 Gated DeltaNet support (b9296 or newer).
# Text chat
llama-cli -m Ornstein3.8-27B-Q4_K_M.gguf -cnv
# Vision
llama-cli -m Ornstein3.8-27B-Q4_K_M.gguf \
--mmproj mmproj-Ornstein3.8-27B-BF16.gguf \
--image photo.jpg -p "Describe this image."
# Server
llama-server -m Ornstein3.8-27B-Q4_K_M.gguf \
--mmproj mmproj-Ornstein3.8-27B-BF16.gguf \
--host 0.0.0.0 --port 8080 -c 8192
LM Studio, Ollama (via a Modelfile), koboldcpp, and text-generation-webui load these files if their bundled llama.cpp supports Qwen3_5ForConditionalGeneration with Gated DeltaNet.
Reproducing the quants
python llama.cpp/convert_hf_to_gguf.py <model_dir> \
--outtype bf16 --outfile Ornstein3.8-27B-BF16.gguf
python llama.cpp/convert_hf_to_gguf.py <model_dir> \
--mmproj --outtype bf16 \
--outfile mmproj-Ornstein3.8-27B-BF16.gguf
llama-quantize Ornstein3.8-27B-BF16.gguf \
Ornstein3.8-27B-Q4_K_M.gguf Q4_K_M
License
Apache 2.0, inherited from the Qwen 3.8 base release.
- Downloads last month
- 204
4-bit
6-bit
8-bit
