llama3.2-1b-CoT GGUF

GGUF quantizations of harshwardhanjadhav/llama3.2-1b-CoT.

Fine-tuned version of unsloth/Llama-3.2-1B-Instruct on 10,000 samples from the open-thoughts/OpenThoughts-114k dataset.
Trained to produce structured Chain-of-Thought reasoning traces using <|begin_of_thought|> and <|begin_of_solution|> blocks before delivering final answers.
Specializes in multi-step mathematical and analytical problem solving.


Quantization Table

File Bits Size Use Case
model_f16.gguf 16 ~2.5 GB Maximum quality, reference
model_q8_0.gguf 8 ~1.3 GB Best quality / size tradeoff
model_q6_k.gguf 6 ~1.0 GB High quality
model_q5_k_m.gguf 5 ~0.9 GB Recommended
model_q5_k_s.gguf 5 ~0.85 GB Slightly smaller Q5
model_q4_k_m.gguf 4 ~0.75 GB Good balance
model_q4_k_s.gguf 4 ~0.70 GB Smaller Q4
model_q3_k_l.gguf 3 ~0.60 GB Low RAM, large variant
model_q3_k_m.gguf 3 ~0.57 GB Low RAM
model_q3_k_s.gguf 3 ~0.53 GB Minimum RAM Q3
model_q2_k.gguf 2 ~0.45 GB Extreme compression

VRAM Requirements

Quant Min VRAM
F16 4 GB
Q8_0 2 GB
Q4_K_M 1.5 GB
Q2_K 1 GB

System Prompt (required for CoT)

Your role as an assistant involves thoroughly exploring questions through a systematic long thinking process before providing the final precise and accurate solutions.
This requires engaging in a comprehensive cycle of analysis, summarizing, exploration, reassessment, reflection, backtracing, and iteration to develop well-considered thinking process.
Please structure your response into two main sections: Thought and Solution.
In the Thought section, detail your reasoning process using the specified format: <|begin_of_thought|> {thought with steps separated with '\n\n'} <|end_of_thought|>
In the Solution section, present the final solution formatted as: <|begin_of_solution|> {final formatted, precise, and clear solution} <|end_of_solution|>

Recommended Inference Parameters

Parameter Value Reason
temperature 0.1 Stable CoT traces
repetition_penalty 1.15 Prevents infinite reasoning loops
top_p 0.9 Filters low-probability tokens
max_new_tokens 4096 Accommodates long reasoning chains

Usage

llama.cpp

./llama-cli -m model_q4_k_m.gguf \
  --temp 0.1 \
  --repeat-penalty 1.15 \
  --top-p 0.9 \
  -n 4096 \
  -p "A store offers a 20% discount on a $150 coat, then takes 10% more off. What is the final price?"

llama-cpp-python

from llama_cpp import Llama

llm = Llama(model_path="model_q4_k_m.gguf", n_ctx=8192)
output = llm(
    "Solve step by step: 2x + 5 = 13",
    max_tokens=4096,
    temperature=0.1,
    repeat_penalty=1.15,
    top_p=0.9,
)
print(output["choices"][0]["text"])

LM Studio

Search tinyopsec/llama3.2-1b-CoT-GGUF in the model browser.

Ollama

ollama run hf.co/tinyopsec/llama3.2-1b-CoT-GGUF:Q4_K_M

Original Model

harshwardhanjadhav/llama3.2-1b-CoT

Downloads last month
1,485
GGUF
Model size
1B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tinyopsec/llama3.2-1b-CoT-GGUF