Sommerfugl-31B-v2 — GGUF

GGUF builds of Sommerfugl-31B-v2, a Norwegian language model by oolabs.no, for local use with llama.cpp, Ollama and LM Studio.

Built with Gemma. Sommerfugl-31B-v2 is a finetune of google/gemma-4-31B-it, modified by oolabs.no. Use is governed by the Gemma Terms of Use.

Results, known limitations and intended use are on the main model card. The chat template is embedded in every file.

Thinking mode is not supported. Sommerfugl is trained for direct answers; the embedded chat template always runs in non-thinking mode, so no extra flags are needed.

Files

file quant size notes
sommerfugl-31b-v2-Q4_K_M.gguf Q4_K_M ~19 GB default; fits a 24 GB GPU or a 32 GB Mac
sommerfugl-31b-v2-Q5_K_M.gguf Q5_K_M ~22 GB a little closer to full quality
sommerfugl-31b-v2-Q8_0.gguf Q8_0 ~33 GB near full quality

Quantization changes outputs slightly; the benchmark numbers on the main card are for the full-precision weights.

Usage

Ollama

ollama run hf.co/oolabs/sommerfugl-31b-v2-GGUF:Q4_K_M

llama.cpp

llama-server -hf oolabs/sommerfugl-31b-v2-GGUF:Q4_K_M --jinja -ngl 99 -c 8192

LM Studio: search for oolabs/sommerfugl-31b-v2-GGUF and pick a quant.

Downloads last month
82
GGUF
Model size
31B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for oolabs/sommerfugl-31b-v2-GGUF

Quantized
(3)
this model