Ornith-1.5-9B EXL3
EXL3 quantization of ornith-ai/Ornith-1.5-9B for ExLlamaV3 / TabbyAPI.
EXL2 is not possible on this checkpoint. The architecture is Qwen3_5ForConditionalGeneration (hybrid GatedDeltaNet + full attention). Use this repo, not an EXL2 convert.
Converted with ExLlamaV3 on an RTX 3090. MIT, same as the base.
Branches
| Branch | Decoder | lm_head |
Size | Notes |
|---|---|---|---|---|
main / 4.0bpw |
4.0 | 6 | 6.8 G | Default. Fits a 24 GB card with context. |
5.0bpw |
5.0 | 6 | 7.6 G | Extra quality vs 4.0. |
6.0bpw |
6.0 | 6 | 8.4 G | Highest of the three. Two shards. |
Load a branch:
oxfrug/Ornith-1.5-9B-exl3 # 4.0 on main
oxfrug/Ornith-1.5-9B-exl3:4.0bpw
oxfrug/Ornith-1.5-9B-exl3:5.0bpw
oxfrug/Ornith-1.5-9B-exl3:6.0bpw
Convert details
| Tool | ExLlamaV3 convert.py |
| Cal | 250 rows × 2048 cols (library default) |
| Codebook | mul1 |
| Vision | stored unquantized (16-bit) |
| MTP | not included — base config.json sets mtp_num_hidden_layers: 1 but the published safetensors have no mtp.* tensors |
Smoke (not a published SWE/Terminal-Bench run)
On 4.0bpw, greedy-ish ExLlamaV3 Generator, ChatML:
lambda x: x ** 2— correct- Swedish “tre plus fem” → “Tre plus fem är åtta.”
is_even(n)with docstring — correct
Reasoning still emits <think>…</think> as designed. These are not the official BF16 agentic scores from the base card.
Serve
ExLlamaV3 ≥ current master (Qwen 3.5 support). TabbyAPI is the usual OpenAI-compatible front.
Sampling from the base card: coding temperature=0.6, top_p=0.95; general temperature=1.0, presence_penalty=1.5. Tool parser: Qwen3 XML.
License
MIT. Derivative of Ornith-1.5-9B by ornith-ai. Quant by oxfrug.
- Downloads last month
- -
Model tree for oxfrug/Ornith-1.5-9B-exl3
Base model
ornith-ai/Ornith-1.5-9B