GGUF
conversational

OpenAI gpt-oss weights in GGUF format, for Understand's AI features, served locally by ullama (llama.cpp).

  • gpt-oss-20b-Q4_K_M.gguf β€” gpt-oss-20b as quantized by Unsloth, copied unmodified from unsloth/gpt-oss-20b-GGUF (SHA-256 c27536640e410032865dc68781d80a08b98f8db5e93575919af8ccc0568aeb4f). A mixture-of-experts model: 21B parameters, about 3.6B active per token, in an 11.6 GB download. Tested on an Apple M5 MacBook Pro, it qualified for Project Chat (answers match the code 0.663, reads before answering 0.807, follows instructions 0.850) and answered fastest of the qualified models of 4B or more (19 s median). It wrote all 40 test code summaries, with accuracy 0.703 and fact recall 0.500.

These results need low reasoning effort. ullama's ullama-models.conf sets reasoning_effort to low for gpt-oss-20b*. At the model's default (medium), 22 of the 40 summaries spent the whole 4,096-token budget reasoning and returned nothing. If you serve this file with another runtime, set low reasoning effort there.

The larger gpt-oss-120b (63 GB, in two parts) is in SciTools/OnBoard. It was tested at the default medium reasoning, and ullama leaves it there.

Ship and use exactly this file: verdicts do not carry across quantizations.

gpt-oss is released by OpenAI under the Apache 2.0 license and subject to the gpt-oss usage policy. This repository redistributes the weights unmodified apart from quantization; it is not endorsed by OpenAI.

Downloads last month
29
GGUF
Model size
21B params
Architecture
gpt-oss
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for SciTools/gpt-oss

Quantized
(245)
this model