TinyLlama-1.1B-Chat-v1.0-GGUF (Q4_K_M)

Q4_K_M GGUF quantization of TinyLlama/TinyLlama-1.1B-Chat-v1.0, produced with llama.cpp by Quantizelab.dev.

File model-Q4_K_M.gguf
Quantization Q4_K_M
Size on disk 0.67 GB
Base model TinyLlama/TinyLlama-1.1B-Chat-v1.0

Run it (full GPU offload)

GGUF defaults to CPU. To get GPU speed you must offload every layer โ€” a single layer left on the CPU takes generation from ~25 tok/s to ~3 tok/s. -ngl 999 simply means "offload all of them".

# llama.cpp
llama-cli -hf thecodehaider/TinyLlama-1.1B-Chat-v1.0-GGUF:model-Q4_K_M.gguf -ngl 999 -c 4096 -p "Hello"

# local file
llama-cli -m model-Q4_K_M.gguf -ngl 999 -c 4096 -cnv

# OpenAI-compatible server
llama-server -m model-Q4_K_M.gguf -ngl 999 -c 4096 --port 8080
# Ollama
ollama run hf.co/thecodehaider/TinyLlama-1.1B-Chat-v1.0-GGUF:Q4_K_M
# llama-cpp-python
from llama_cpp import Llama
llm = Llama(model_path="model-Q4_K_M.gguf", n_gpu_layers=-1, n_ctx=4096)
print(llm("Hello", max_tokens=128)["choices"][0]["text"])

Will it fit your GPU?

Weights plus ~1.2 GB of KV-cache and compute overhead at a 4k context.

GPU VRAM Fits fully offloaded? Headroom for context
NVIDIA T4 / RTX 4060 16 GB Yes ~14.1 GB (4k+ context)
RTX 3090 / 4090 / A10 24 GB Yes ~22.1 GB (4k+ context)
A100 40GB 40 GB Yes ~38.1 GB (4k+ context)

If a row says No, lower -ngl until it fits, or run on CPU (GGUF works either way โ€” it is just slower).

Notes

  • Q4_K_M is the recommended balance of size and quality; Q8_0 and above will not fully offload to a 16 GB card for models past ~8B.
  • Reduce -c (context) first when you hit out-of-memory: the KV cache grows linearly with context length.
Downloads last month
196
GGUF
Model size
1B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for thecodehaider/TinyLlama-1.1B-Chat-v1.0-GGUF

Quantized
(150)
this model