llama-3.2-1b-instruct-fine-tuned GGUF

GGUF quantizations of ai-nexuz/llama-3.2-1b-instruct-fine-tuned.

Fine-tuned version of meta-llama/Llama-3.2-1B-Instruct on the kanhatakeyama/wizardlm8x22b-logical-math-coding-sft dataset.
Specializes in logical reasoning, mathematics, and code generation.


Quantization Table

File Bits Size Use Case
model_f16.gguf 16 ~2.5 GB Maximum quality, reference
model_q8_0.gguf 8 ~1.3 GB Best quality / size tradeoff
model_q6_k.gguf 6 ~1.0 GB High quality
model_q5_k_m.gguf 5 ~0.9 GB Recommended
model_q5_k_s.gguf 5 ~0.85 GB Slightly smaller Q5
model_q4_k_m.gguf 4 ~0.75 GB Good balance
model_q4_k_s.gguf 4 ~0.70 GB Smaller Q4
model_q3_k_l.gguf 3 ~0.60 GB Low RAM, large variant
model_q3_k_m.gguf 3 ~0.57 GB Low RAM
model_q3_k_s.gguf 3 ~0.53 GB Minimum RAM Q3
model_q2_k.gguf 2 ~0.45 GB Extreme compression

VRAM Requirements

Quant Min VRAM
F16 4 GB
Q8_0 2 GB
Q4_K_M 1.5 GB
Q2_K 1 GB

Usage

llama.cpp

./llama-cli -m model_q4_k_m.gguf -p "Solve: 2x + 5 = 13" -n 256

llama-cpp-python

from llama_cpp import Llama
llm = Llama(model_path="model_q4_k_m.gguf", n_ctx=2048)
output = llm("Solve: 2x + 5 = 13", max_tokens=256)
print(output["choices"][0]["text"])

LM Studio

Search tinyopsec/llama-3.2-1b-instruct-fine-tuned-GGUF in the model browser.

Ollama

ollama run hf.co/tinyopsec/llama-3.2-1b-instruct-fine-tuned-GGUF:Q4_K_M

Original Model

ai-nexuz/llama-3.2-1b-instruct-fine-tuned

Downloads last month
1,478
GGUF
Model size
1B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tinyopsec/llama-3.2-1b-instruct-fine-tuned-GGUF