Ornith-1.5-9B EXL3 4.0 bpw HQ

An EXL3 conversion of ornith-ai/Ornith-1.5-9B, converted with ExLlamaV3 1.4.2. This is a community-made quantization by ultimatechris, not an official Ornith release.

Quantization

  • Converter: ExLlamaV3 1.4.2
  • Recipe: high-quality (-hq) EXL3 conversion
  • Target bitrate: 4.0 bpw
  • Conversion hardware: one NVIDIA A100 40GB

The HQ pass uses the target bitrate for the decoder and keeps selected small or sensitive tensors at higher precision. The language-model head is stored at 6 bpw. The exact tensor layout is recorded in quantization_config.json.

The artifact loads and generates successfully with ExLlamaV3 1.4.2. A short deterministic smoke test produced 31 tokens at 20.9 tokens/second on an NVIDIA A100-40GB. This is a load-and-generation check rather than a broad quality benchmark.

Base-vs-quantization check

A small teacher-forced comparison scored the same 60 target tokens from five fixed text continuations with the original BF16 base and this EXL3 artifact:

  • Renormalized KL on the shared finite vocabulary support: 0.01375
  • Next-token top-1 agreement: 93.3%
  • NLL delta (EXL3 minus base): -0.0431 nats/token

No meaningful degradation was observed in this sample. This is a limited diagnostic, not a broad benchmark; the slightly negative NLL delta should not be interpreted as evidence that the quantization improves the base model.

Runtime

Use an ExLlamaV3-compatible frontend such as TabbyAPI. Point the frontend at the repository root. The repository includes the tokenizer, chat template, configuration, and processor files needed by the base model.

Runtime speed and maximum context depend on the frontend, GPU, context length, and batch settings.

Attribution and limitations

The base model was developed and released by the Ornith team. This quantized file retains the base model's capabilities, limitations, and usage requirements. Read the base model card before deployment.

Downloads last month
257
Safetensors
Model size
4B params
Tensor type
BF16
·
F16
·
I16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ultimatechris/Ornith-1.5-9B-EXL3-4bpw

Quantized
(37)
this model