Astrea R8 Chat 9B — Q8_0 GGUF

This is an unofficial community Q8_0 GGUF conversion of Altworld/Astrea-R8-Chat-9B for llama.cpp, with an explicit non-reasoning chat template.

File

File Quantization Size
Astrea-R8-Chat-9B-Q8_0.gguf Q8_0 9,527,501,280 bytes (8.87 GiB)

Embedded non-reasoning template

The GGUF contains a hard non-reasoning Jinja template in tokenizer.chat_template; no external template file is required.

Astrea's optional thinking mode was not reliable in local llama.cpp testing: simple prompts could consume hundreds of tokens before emitting </think>, and often did not end reasoning at all, and simply responded as if reasoning was not enabled. The bundled template therefore always places a closed, empty thinking block in the prompt and does not expose an enable_thinking template variable. It also omits hidden reasoning when replaying assistant messages into conversation history. A standalone copy is included as chat_template.jinja for inspection.

llama.cpp

llama-server.exe `
  --model Astrea-R8-Chat-9B-Q8_0.gguf `
  --jinja `
  --reasoning off `
  --reasoning-format none `
  --ctx-size 32768 `
  --n-gpu-layers all `
  --temp 0.8 `
  --top-p 1.0 `
  --top-k 0 `
  --min-p 0.025 `
  --repeat-penalty 1.08

The model metadata advertises a 262,144-token context window. Choose a context size appropriate for your available VRAM/RAM. The command above starts at a more conservative 32,768 tokens. I was able to easily run a much more ambitous setup with -ngl all --fit off -c 147456 -np 4 --kv-unified on a 16GB VRAM card (5070 Ti).

Conversion notes

  1. The original safetensors were converted to BF16 GGUF with llama.cpp's convert_hf_to_gguf.py using --no-mtp. The downloaded checkpoint did not contain the extra MTP-layer tensors declared by its configuration.
  2. BF16 was quantized with llama-quantize using Q8_0.
  3. llama.cpp's gguf_new_metadata.py embedded the hard non-reasoning template; this metadata-only copy did not requantize tensors.

The tensor-only SHA-256 reported by llama-gguf-hash was identical before and after the metadata rewrite:

20d213a0c5ee663ef6d02ffcff8d0b28cbff18b559c67aca5250cd5e6a22d624

The final whole-file checksums are in SHA256SUMS.

Validation

The final GGUF was loaded directly by llama-server without --chat-template-file. Its exposed template matched the bundled standalone Jinja, and a request that explicitly supplied enable_thinking=true still returned normal content with no reasoning_content.

License and attribution

The source model is released under Apache-2.0. See LICENSE and NOTICE, and refer to the source model card for its intended use, evaluation results, and limitations.

Downloads last month
159
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Michionlion/Astrea-R8-Chat-9B-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(4)
this model