Nanbeige4.2-3B-Q4_K_M

Q4_K_M GGUF quantization of Nanbeige/Nanbeige4.2-3B.

Details

  • Architecture: nanbeige
  • Format: GGUF
  • Quantization: Q4_K_M
  • Effective BPW: 4.93
  • Original size: 7953.53 MiB
  • Quantized size: 2451.74 MiB
  • Runtime: llama.cpp

Usage

Requires a recent llama.cpp version with nanbeige support.

llama-cli -m Nanbeige4.2-3B-Q4_K_M.gguf -ngl 99

Evaluation

Wikitext-2:

  • Perplexity: 26.8796 ± 0.2620
  • Context: 512
  • Batch size: 512
  • GPU offload: 99 layers

Base Model

Nanbeige/Nanbeige4.2-3B

No additional training or fine-tuning was performed.

Downloads last month
36
GGUF
Model size
4B params
Architecture
nanbeige
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mackkkkkilllll/Nanbeige4.2-3B-Q4_K_M

Quantized
(45)
this model