Vikhr-Llama-3.2-1B-Instruct GGUF

GGUF quantizations of Vikhrmodels/Vikhr-Llama-3.2-1B-Instruct.

Quantization Table

File Bits Size Use Case
model_f16.gguf 16 ~2.5 GB Maximum quality, reference
model_q8_0.gguf 8 ~1.3 GB Best quality, near-lossless
model_q6_k.gguf 6 ~1.0 GB Great quality
model_q5_k_m.gguf 5 ~0.9 GB Balanced quality/size
model_q5_k_s.gguf 5 ~0.85 GB Slightly smaller Q5
model_q4_k_m.gguf 4 ~0.77 GB Recommended, good balance
model_q4_k_s.gguf 4 ~0.72 GB Smaller Q4
model_q3_k_l.gguf 3 ~0.65 GB Low VRAM, decent quality
model_q3_k_m.gguf 3 ~0.60 GB Low VRAM
model_q3_k_s.gguf 3 ~0.55 GB Very low VRAM
model_q2_k.gguf 2 ~0.45 GB Minimum size, lowest quality

VRAM Requirements

Quant VRAM
F16 ~3 GB
Q8_0 ~1.8 GB
Q6_K ~1.5 GB
Q5_K_M ~1.3 GB
Q4_K_M ~1.1 GB
Q3_K_M ~0.8 GB
Q2_K ~0.6 GB

Usage

llama.cpp

./llama-cli -m model_q4_k_m.gguf -p "Your prompt here" -n 512

llama-cpp-python

from llama_cpp import Llama

llm = Llama(
    model_path="model_q4_k_m.gguf",
    n_ctx=2048,
    n_gpu_layers=-1
)

response = llm(
    "Your prompt here",
    max_tokens=512,
    stop=["<|eot_id|>"]
)
print(response["choices"][0]["text"])

LM Studio

Download any .gguf file and load it directly in LM Studio.

Ollama

ollama run hf.co/tinyopsec/Vikhr-Llama-3.2-1B-Instruct-GGUF

Original Model

Vikhrmodels/Vikhr-Llama-3.2-1B-Instruct

Quantized using llama.cpp.

Downloads last month
1,758
GGUF
Model size
1B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for tinyopsec/Vikhr-Llama-3.2-1B-Instruct-GGUF