Instructions to use tinyopsec/Skywork-OR1-7B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use tinyopsec/Skywork-OR1-7B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf tinyopsec/Skywork-OR1-7B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf tinyopsec/Skywork-OR1-7B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf tinyopsec/Skywork-OR1-7B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf tinyopsec/Skywork-OR1-7B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf tinyopsec/Skywork-OR1-7B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf tinyopsec/Skywork-OR1-7B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf tinyopsec/Skywork-OR1-7B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf tinyopsec/Skywork-OR1-7B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/tinyopsec/Skywork-OR1-7B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use tinyopsec/Skywork-OR1-7B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tinyopsec/Skywork-OR1-7B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tinyopsec/Skywork-OR1-7B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/tinyopsec/Skywork-OR1-7B-GGUF:Q4_K_M
- Ollama
How to use tinyopsec/Skywork-OR1-7B-GGUF with Ollama:
ollama run hf.co/tinyopsec/Skywork-OR1-7B-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use tinyopsec/Skywork-OR1-7B-GGUF with Docker Model Runner:
docker model run hf.co/tinyopsec/Skywork-OR1-7B-GGUF:Q4_K_M
- Lemonade
How to use tinyopsec/Skywork-OR1-7B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull tinyopsec/Skywork-OR1-7B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Skywork-OR1-7B-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Skywork-OR1-7B GGUF
GGUF quantizations of Skywork/Skywork-OR1-7B.
Description
Skywork-OR1-7B (Open Reasoner 1) is a general-purpose math and code reasoning model trained with large-scale rule-based reinforcement learning using a customized GRPO algorithm. It is based on DeepSeek-R1-Distill-Qwen-7B and trained on 110K math problems and 14K coding questions with model-aware difficulty estimation, offline/online filtering, rejection sampling, multi-stage training pipeline, and adaptive entropy control.
Quantization Files
| File | Bits | Size | Use Case |
|---|---|---|---|
| model_f16.gguf | 16 | ~15.2 GB | Maximum quality, reference |
| model_q8_0.gguf | 8 | ~8.5 GB | Best quality, near-lossless |
| model_q6_k.gguf | 6 | ~6.6 GB | High quality |
| model_q5_k_m.gguf | 5 | ~5.7 GB | Balanced quality |
| model_q5_k_s.gguf | 5 | ~5.5 GB | Balanced quality, smaller |
| model_q4_k_m.gguf | 4 | ~4.9 GB | Good quality, recommended |
| model_q4_k_s.gguf | 4 | ~4.7 GB | Good quality, smaller |
| model_q3_k_l.gguf | 3 | ~4.0 GB | Low quality, small |
| model_q3_k_m.gguf | 3 | ~3.7 GB | Low quality, smaller |
| model_q3_k_s.gguf | 3 | ~3.5 GB | Very low quality |
| model_q2_k.gguf | 2 | ~3.0 GB | Lowest quality, minimum size |
VRAM Requirements
| Quantization | VRAM |
|---|---|
| F16 | ~16 GB |
| Q8_0 | ~9 GB |
| Q6_K | ~7 GB |
| Q5_K_M | ~6 GB |
| Q4_K_M | ~5 GB |
| Q3_K_M | ~4 GB |
| Q2_K | ~3.5 GB |
Usage
llama.cpp
./llama-cli -m model_q4_k_m.gguf -p "<your prompt>" -n 512
llama-cpp-python
from llama_cpp import Llama
llm = Llama(model_path="model_q4_k_m.gguf", n_ctx=32768)
output = llm("<your prompt>", max_tokens=512)
print(output["choices"][0]["text"])
LM Studio
Load any .gguf file directly in LM Studio via Load Model.
Ollama
ollama run hf.co/tinyopsec/Skywork-OR1-7B-GGUF:Q4_K_M
Original Model
- Skywork/Skywork-OR1-7B
- Training data: Skywork/Skywork-OR1-RL-Data
- License: Apache 2.0
- Downloads last month
- 1,271
2-bit
3-bit
4-bit
5-bit
6-bit
8-bit
16-bit
Model tree for tinyopsec/Skywork-OR1-7B-GGUF
Base model
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B