NB-Llama-3.1-8B-Instruct โ€” EXL2

ExLlamaV2 / EXL2 quants of NbAiLab/nb-llama-3.1-8B-Instruct.

Official Hub files are BF16 and GGUF. There was no EXL2 pack.

Converted with ExLlamaV2 0.3.2, lm_head at 6-bit, built-in default calibration. One measurement pass, then each bitrate from measurement.json.

Branches

Revision Target bpw
4.0bpw 4.0
4.5bpw 4.5
5.0bpw 5.0
5.5bpw 5.5
6.0bpw 6.0
hf download oxfrug/nb-llama-3.1-8B-Instruct-exl2 --revision 5.0bpw --local-dir ./nb-llama-3.1-8B-Instruct-exl2-5.0bpw

Notes

  • License: Meta Llama 3.1 Community License. Keep NOTICE.
  • Loader: ExLlamaV2. Not GGUF.
  • On PyTorch 2.13 without Flash Attention 2.5.7+, set config.no_sdpa = True before load.

Source

NbAiLab/nb-llama-3.1-8B-Instruct
  โ† meta-llama/Llama-3.1-8B-Instruct
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for oxfrug/nb-llama-3.1-8B-Instruct-exl2