Skywork-OR1-7B GGUF

GGUF quantizations of Skywork/Skywork-OR1-7B.

Description

Skywork-OR1-7B (Open Reasoner 1) is a general-purpose math and code reasoning model trained with large-scale rule-based reinforcement learning using a customized GRPO algorithm. It is based on DeepSeek-R1-Distill-Qwen-7B and trained on 110K math problems and 14K coding questions with model-aware difficulty estimation, offline/online filtering, rejection sampling, multi-stage training pipeline, and adaptive entropy control.

Quantization Files

File Bits Size Use Case
model_f16.gguf 16 ~15.2 GB Maximum quality, reference
model_q8_0.gguf 8 ~8.5 GB Best quality, near-lossless
model_q6_k.gguf 6 ~6.6 GB High quality
model_q5_k_m.gguf 5 ~5.7 GB Balanced quality
model_q5_k_s.gguf 5 ~5.5 GB Balanced quality, smaller
model_q4_k_m.gguf 4 ~4.9 GB Good quality, recommended
model_q4_k_s.gguf 4 ~4.7 GB Good quality, smaller
model_q3_k_l.gguf 3 ~4.0 GB Low quality, small
model_q3_k_m.gguf 3 ~3.7 GB Low quality, smaller
model_q3_k_s.gguf 3 ~3.5 GB Very low quality
model_q2_k.gguf 2 ~3.0 GB Lowest quality, minimum size

VRAM Requirements

Quantization VRAM
F16 ~16 GB
Q8_0 ~9 GB
Q6_K ~7 GB
Q5_K_M ~6 GB
Q4_K_M ~5 GB
Q3_K_M ~4 GB
Q2_K ~3.5 GB

Usage

llama.cpp

./llama-cli -m model_q4_k_m.gguf -p "<your prompt>" -n 512

llama-cpp-python

from llama_cpp import Llama

llm = Llama(model_path="model_q4_k_m.gguf", n_ctx=32768)
output = llm("<your prompt>", max_tokens=512)
print(output["choices"][0]["text"])

LM Studio

Load any .gguf file directly in LM Studio via Load Model.

Ollama

ollama run hf.co/tinyopsec/Skywork-OR1-7B-GGUF:Q4_K_M

Original Model

Downloads last month
1,271
GGUF
Model size
8B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tinyopsec/Skywork-OR1-7B-GGUF

Quantized
(6)
this model