Ornith-1.5-9B EXL3

EXL3 quantization of ornith-ai/Ornith-1.5-9B for ExLlamaV3 / TabbyAPI.

EXL2 is not possible on this checkpoint. The architecture is Qwen3_5ForConditionalGeneration (hybrid GatedDeltaNet + full attention). Use this repo, not an EXL2 convert.

Converted with ExLlamaV3 on an RTX 3090. MIT, same as the base.

Branches

Branch Decoder lm_head Size Notes
main / 4.0bpw 4.0 6 6.8 G Default. Fits a 24 GB card with context.
5.0bpw 5.0 6 7.6 G Extra quality vs 4.0.
6.0bpw 6.0 6 8.4 G Highest of the three. Two shards.

Load a branch:

oxfrug/Ornith-1.5-9B-exl3          # 4.0 on main
oxfrug/Ornith-1.5-9B-exl3:4.0bpw
oxfrug/Ornith-1.5-9B-exl3:5.0bpw
oxfrug/Ornith-1.5-9B-exl3:6.0bpw

Convert details

Tool ExLlamaV3 convert.py
Cal 250 rows × 2048 cols (library default)
Codebook mul1
Vision stored unquantized (16-bit)
MTP not included — base config.json sets mtp_num_hidden_layers: 1 but the published safetensors have no mtp.* tensors

Smoke (not a published SWE/Terminal-Bench run)

On 4.0bpw, greedy-ish ExLlamaV3 Generator, ChatML:

  • lambda x: x ** 2 — correct
  • Swedish “tre plus fem” → “Tre plus fem är åtta.”
  • is_even(n) with docstring — correct

Reasoning still emits <think>…</think> as designed. These are not the official BF16 agentic scores from the base card.

Serve

ExLlamaV3 ≥ current master (Qwen 3.5 support). TabbyAPI is the usual OpenAI-compatible front.

Sampling from the base card: coding temperature=0.6, top_p=0.95; general temperature=1.0, presence_penalty=1.5. Tool parser: Qwen3 XML.

License

MIT. Derivative of Ornith-1.5-9B by ornith-ai. Quant by oxfrug.

Downloads last month
-
Safetensors
Model size
4B params
Tensor type
BF16
·
F16
·
I16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for oxfrug/Ornith-1.5-9B-exl3

Quantized
(15)
this model