Instructions to use prometheusAIR/Solar-Open2-250B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use prometheusAIR/Solar-Open2-250B-GGUF with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="prometheusAIR/Solar-Open2-250B-GGUF", filename="IQ1_M/Solar-Open2-250B-IQ1_M-00001-of-00002.gguf", )
llm.create_chat_completion( messages = [ { "role": "user", "content": "What is the capital of France?" } ] ) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use prometheusAIR/Solar-Open2-250B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M # Run inference directly in the terminal: llama cli -hf prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M # Run inference directly in the terminal: llama cli -hf prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M # Run inference directly in the terminal: ./llama-cli -hf prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M
Use Docker
docker model run hf.co/prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M
- LM Studio
- Jan
- vLLM
How to use prometheusAIR/Solar-Open2-250B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prometheusAIR/Solar-Open2-250B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prometheusAIR/Solar-Open2-250B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M
- Ollama
How to use prometheusAIR/Solar-Open2-250B-GGUF with Ollama:
ollama run hf.co/prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M
- Unsloth Studio
How to use prometheusAIR/Solar-Open2-250B-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for prometheusAIR/Solar-Open2-250B-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for prometheusAIR/Solar-Open2-250B-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for prometheusAIR/Solar-Open2-250B-GGUF to start chatting
- Pi
How to use prometheusAIR/Solar-Open2-250B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use prometheusAIR/Solar-Open2-250B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use prometheusAIR/Solar-Open2-250B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use prometheusAIR/Solar-Open2-250B-GGUF with Docker Model Runner:
docker model run hf.co/prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M
- Lemonade
How to use prometheusAIR/Solar-Open2-250B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull prometheusAIR/Solar-Open2-250B-GGUF:IQ1_M
Run and chat with the model
lemonade run user.Solar-Open2-250B-GGUF-IQ1_M
List all available models
lemonade list
llm.create_chat_completion(
messages = [
{
"role": "user",
"content": "What is the capital of France?"
}
]
)Solar-Open2-250B-GGUF
GGUF quantizations of upstage/Solar-Open2-250B — Upstage's 250B-A15B hybrid linear-attention MoE — for local inference with llama.cpp. Built and tested on a workstation with 128 GB DDR5 system RAM and an RTX 6000 Pro 96 GB Blackwell GPU.
⚠️ These files do NOT run on mainline llama.cpp (yet)
Solar Open 2 uses a new architecture (
solar_open2) that is not in upstream llama.cpp. Loading these GGUFs with a stock build fails withunknown model architecture: 'solar-open2'.You must build the small patch included here (
solar_open2-llama.cpp.patch, ~775 lines across 10 files). See Building llama.cpp below. If/when the architecture is merged upstream, this requirement goes away.
Built with Solar.
This is an independent, third-party quantization. It is not affiliated with or endorsed by Upstage. All model weights are © Upstage under the Upstage Solar License.
What is Solar Open 2?
A 250B-parameter Mixture-of-Experts that activates ~15B parameters per token. Its stack interleaves one softmax-attention (GQA) layer with three linear-attention (KDA) layers in every block of four (S-L-L-L × 12 = 48 layers), with no positional encoding (NoPE). The linear layers carry sequence context in a fixed-size recurrent state, so the KV cache stays constant in length and only the 12 softmax layers cache anything — making very long context cheap. Trained on English, Korean, and Japanese, aimed at agentic / office / coding use. See the technical report for details.
The architecture's linear-attention core is KDA (Kimi Delta Attention) with a negative-eigenvalue extension (β = 2σ(·) ∈ (0,2)), the same delta-net family already supported in llama.cpp for Kimi-Linear; this port reuses that kernel and adds the GQA-with-sigmoid-gate softmax path, NoPE, and the S-L-L-L layout.
Available quantizations
| File | Bits | Size | PPL (wikitext-2)¹ | vs Q6_K |
|---|---|---|---|---|
Q6_K/ |
6.57 | 191 GiB | 4.123 | — |
IQ4_XS/ ⭐ |
4.35 | 127 GiB | 4.159 | +0.9 % |
Q2_K/ |
3.05 | 89 GiB | 4.774 | +15.8 % |
IQ1_M/ ⚠️ |
1.93 | 56 GiB | 7.717 | +87.2 % |
¹ Perplexity over a fixed 40-chunk slice of wiki.test.raw at n_ctx=512. Lower is better; all four were measured identically so the deltas are directly comparable. Q6_K is the near-lossless baseline the deltas are computed against.
On comparing perplexity across repos: these numbers are only meaningful against an identical slice. Our own
Q2_Kmeasures 4.774 over 40 chunks but 5.486 over 80 chunks — same file, same settings, different chunk count. If you are comparing our numbers to another quantizer's, check that the chunk count matches first.
Which to pick:
IQ4_XS— recommended. Near-Q6_K quality (+0.88 % PPL is within measurement noise) at two-thirds the size. This is the quality/size sweet spot.Q2_K— speed / single-card. Fits entirely on one 96 GB GPU and runs ~2.5× faster, but +15.8 % PPL is a real, perceptible quality drop (the sub-4-bit cliff on this fine-grained 320-expert MoE). Good for drafts, high-throughput, or demos; not a replacement for IQ4_XS as a daily driver.Q6_K— near-lossless reference. For evaluation or as an imatrix source.IQ1_M— smaller-VRAM only. Read the warning below before choosing this version.
⚠️ About IQ1_M
Perplexity is 87 % worse than Q6_K That is a severe, honest number and we are putting it up front rather than in a footnote. It was built by request, for people who cannot fit the 89 GiB Q2_K.
IQ1_M gains are mostly for memory footprint, not speed. Measured on one RTX PRO 6000 (96 GB), both fully GPU-resident: IQ1_M decodes at 95.5 tok/s using 58 GiB VRAM, Q2_K at 94.9 tok/s using ~89 GiB. Identical speed. If you can fit Q2_K, use Q2_K — IQ1_M is intended for 64 GB-class hardware (64 GB Apple Silicon, 2×32 GB, 80 GB A100).
That said, it degrades more gracefully than the perplexity implies. At temperature 0 it answered factual chains, arithmetic (47 × 23 = 1081, self-verified), a correct Rayleigh-scattering explanation, a correct recursive fib, and fluent Korean translation — and it scored 15/15 exact on needle-in-a-haystack at 2K/4K/8K/16K/28K × 10 %/50 %/90 % depth. Our reading is that the 320-expert MoE has enough routing redundancy, and keeping attention and embeddings at Q6_K preserves enough of the retrieval machinery, that the damage shows up as degraded fluency and reliability rather than outright breakage. Expect more repetition, weaker long-form coherence, and more frequent factual slips than any other rung here.
All four protect what matters most: attention and embeddings are kept at Q6_K (only ~3 % of parameters, so the cost is negligible), and blk.0's experts — which the calibration set activates least — are boosted (Q5_K in IQ4_XS, Q4_K in Q2_K, IQ2_XXS in IQ1_M) to avoid quantizing rarely-seen weights against a weak importance signal. Quantization used the included importance matrix (Solar-Open2-250B.imatrix), computed over an English + Korean + Japanese + code calibration mix so language- and domain-specialized experts are represented.
Blackwell / sm_120 note. The
iq1_s,iq2_s, andiq3_sCUDA matmul kernels produce wrong results on compute capability 12.0 (verified withtest-backend-opson both an RTX PRO 6000 and a 5070 Ti). That is an upstream llama.cpp issue, not specific to this model — but it means any low-bit quant built around those types (e.g.IQ2_M, which usesiq2_s) is unusable on 50-series and RTX PRO Blackwell cards. The types used here —iq1_m,iq2_xxs,iq4_xs,iq4_nl, and all K-quants — all pass.
Building llama.cpp
The patch applies on top of upstream commit 846e991ec (any nearby master should work with minor fuzz).
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
git checkout 846e991ec
git apply /path/to/solar_open2-llama.cpp.patch # adds the solar_open2 arch + tool parser
cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=120 # 120 = Blackwell; set to your GPU
cmake --build build --target llama-server -j
The patch touches: gguf-py + src/llama-arch.* (arch registration), conversion/solar_open2.py + convert_hf_to_gguf.py wiring (HF→GGUF conversion), src/models/solar-open2.cpp (the graph), common/chat.cpp (tool-call + reasoning parser), and src/llama-vocab.cpp (registers <|im:end|> as an end-of-generation token).
Note: if you built from an earlier version of this patch and saw
<|im:end|>at the end of replies, rebuild with the current patch — the fix is here, and your existing GGUF downloads work as-is (no re-download needed).
Running
Download all shards of a quant into one directory, then point --model at the first shard (…-00001-of-000NN.gguf); llama.cpp loads the rest automatically.
./build/bin/llama-server \
--model IQ4_XS/Solar-Open2-250B-IQ4_XS-00001-of-00004.gguf \
--host 0.0.0.0 --port 8080 \
--device CUDA0 -ngl 999 -ncmoe 15 \
--no-mmap \
--ctx-size 32768 -fa on -b 512 -ub 512 -np 1 --threads 12 \
--jinja
Flags that are load-bearing, not decoration:
--no-mmapis required for CPU-offloaded setups. With mmap, the CPU-resident experts are demand-faulted from disk in 4 KB pages and throughput collapses (1.9 tok/s in our tests);18 tok/s at IQ4_XS with offload). Not needed if the whole model is on the GPU.--no-mmaploads them into RAM up front (-ub 512(or4096), never2048. This port has a bug at-ub 2048specifically: a full 2048-token micro-batch followed by a small tail triggers a CUDA illegal-memory-access on this architecture's shape.-ub 512and-ub 4096are both fine. (Reported below under Known issues.)-ncmoe Nand--device CUDA0.-ncmoekeeps the experts of the first N layers on CPU. Use a single GPU device with it — llama.cpp assigns later (expert-heavy) layers to additional devices, which will OOM a smaller second GPU. Drop-ncmoeentirely if the model fits on one GPU (e.g. Q2_K on a 96 GB card).
Rough hardware sizing (per quant, for -ncmoe tuning): weights split between VRAM and system RAM. IQ4_XS ≈ 127 GiB total (e.g. ~92 GiB VRAM + ~35 GiB RAM at -ncmoe 15); Q2_K ≈ 89 GiB fits on a single 96 GB card with no offload; IQ1_M ≈ 56 GiB fits a 64 GB card or 64 GB of unified memory with no offload. The KV cache is tiny — only 12 of 48 layers cache, ~48 KiB/token — so long context costs little VRAM (256K ≈ 12 GiB, 1M ≈ 48 GiB f16).
Budgeting VRAM at very long context. The KV cache is not the only term that grows. At -ub 4096 and a 1M context the CUDA compute buffer alone is ~12.9 GiB (it scales with -ub, and is only ~2 GiB at 128K). A working budget from our own tuning, per GPU: non-expert weights + (2.55 GiB × resident expert layers) + KV + compute buffer + ~2.7 GiB overhead. Size the compute buffer before you size the offload, or you will OOM after the weights have already loaded.
Measured performance
On a single RTX PRO 6000 Blackwell (96 GB) + Ryzen 9950X3D, DDR5:
| Quant | Offload | Decode | Prefill | Needle retrieval |
|---|---|---|---|---|
| IQ4_XS | -ncmoe 15 |
37.6 tok/s | 31.8 tok/s | exact to 16K (tested) |
| Q2_K | all-GPU | 94.9 tok/s | 63.5 tok/s | exact to 4K (tested) |
| IQ1_M | all-GPU (58 GiB) | 95.5 tok/s | — | 15/15 exact to 28K |
Reasoning
Solar Open 2 is a reasoning model, and reasoning is ON by default — the template's first line is set reasoning_effort = reasoning_effort if reasoning_effort is defined else "high", so omitting the setting gives you full-depth thinking. To control it, pass reasoning_effort in chat_template_kwargs:
{
"model": "solar",
"messages": [ ... ],
"chat_template_kwargs": { "reasoning_effort": "high" }
}
Values that turn reasoning on: medium, high, xhigh. Any other value (e.g. "none") turns it off — but omitting the key entirely falls back to the template's own default of "high", i.e. on.
Agent/client note. Because the default is on, clients that can't set
chat_template_kwargsget full-depth reasoning on every call — including cheap auxiliary ones like session-title generation. Measured here: a title request reasoned for 3,743 tokens (~105 s) to produce a four-word title, versus 5 tokens withreasoning_effort: "none". If your client has short timeouts on such helper calls, either raise them or bound thinking server-side with--reasoning-budget N(note that this caps main-task reasoning too).
Tool calling
Supported via the standard OpenAI tools / tool_calls API. The included parser handles Solar's custom marker format (<|tool_call:start|>name … <|tool_arg:start|>key<|tool_arg:value|>value<|tool_arg:end|> … <|tool_call:end|>), including multiple/parallel calls and typed arguments. Start the server with --jinja (shown above) and send tools as usual.
Provenance & reproducibility
- Base weights:
upstage/Solar-Open2-250B(BF16), converted to a BF16 GGUF, then quantized. - Importance matrix: computed from the Q6_K over an English (
calibration_datav3) + Korean/Japanese Wikipedia + source-code mix, so that language- and domain-specialized experts are exercised. Included asSolar-Open2-250B.imatrix(GGUF format) if you want to make your own quants. - Per-tensor recipe (both builds): experts at the headline type;
blk.0experts boosted one K-quant rung (weakest calibration coverage); allattn_*and token/output embeddings at Q6_K. - Every build was validated for coherence, needle-in-haystack retrieval at depth, and perplexity before release.
IQ4_XS's numeric correctness was additionally checked against an independent NumPy reference forward pass of the architecture.
Known issues
-ub 2048crashes (CUDA illegal memory access) on this architecture's shape when a full 2048-token micro-batch is followed by a small tail. Use-ub 512or-ub 4096. Under investigation; the weights/conversion are not implicated (a runtime flag fully avoids it).- This is a community port, not upstream-reviewed. The architecture, quant recipe, reasoning behavior, and tool parser are validated empirically, but edge cases may remain. Issues/feedback welcome.
Attribution & license
Model weights are © Upstage, released under the Upstage Solar License (see LICENSE). Per that license, this derivative is named with the "Solar" prefix, displays "Built with Solar," and ships a copy of the license. The quantization tooling and the llama.cpp patch are provided under llama.cpp's MIT license.
If you use this, please cite the Solar Open 2 technical report and link back to the original model.
- Downloads last month
- 3,258
1-bit
2-bit
4-bit
6-bit
Model tree for prometheusAIR/Solar-Open2-250B-GGUF
Base model
upstage/Solar-Open2-250B
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="prometheusAIR/Solar-Open2-250B-GGUF", filename="", )