granite-4.2-8B GGUF

GGUF conversions of ibm-granite/granite-4.2-8b for llama.cpp at 16-bit, 8-bit and 4-bit precision.

Files

File Precision Size
granite-4.2-8b-BF16.gguf BF16 (16-bit, lossless from source) 17.6 GB
granite-4.2-8b-Q8_0.gguf Q8_0 (8-bit, 8.50 bits per weight) 9.3 GB
granite-4.2-8b-Q4_K_M.gguf Q4_K_M (4-bit, 4.86 bits per weight) 5.3 GB

Usage

# Chat server with Granite's embedded chat template (tool calling and thinking)
llama-server -hf webAI-Official/granite-4.2-8B:Q4_K_M --jinja

# Interactive chat
llama-cli -hf webAI-Official/granite-4.2-8B:Q8_0

Thinking is on by default in the chat template. Pass --chat-template-kwargs '{"enable_thinking": false}' to llama-server to turn it off.

How these files were made

  1. The source safetensors (bf16) were converted with llama.cpp b9770: convert_hf_to_gguf.py ibm-granite/granite-4.2-8b --outtype bf16
  2. Q8_0 and Q4_K_M were each quantized directly from the BF16 file with llama-quantize. No importance matrix was used.

Performance

Measured with llama-bench (llama.cpp b9770, Metal) on an Apple M5 Pro with 24 GB of memory:

File Prompt processing, 512 tokens (tok/s) Generation, 128 tokens (tok/s)
BF16 834 15.3
Q8_0 1080 29.1
Q4_K_M 1051 47.8

On a 24 GB machine, BF16 does not fit in GPU memory with llama.cpp's default context size. Use a smaller context, such as -c 2048, or choose Q8_0.

License

Apache 2.0, the same as the base model. See ibm-granite/granite-4.2-8b for the model's details, intended use and limitations.

Downloads last month
109
GGUF
Model size
9B params
Architecture
granite
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for webAI-Official/granite-4.2-8B

Quantized
(47)
this model