Qwen3-4B GGUF โ€” ready-to-run quantized builds

Pre-quantized GGUF builds of Qwen/Qwen3-4B for use with llama.cpp, Ollama, LM Studio, and any GGUF-compatible runtime.

We maintain these so you can run a strong 4B model locally in minutes โ€” and when you need 340+ frontier models instead, our API at qubax.ai serves them at some of the lowest prices on the market (pay with crypto, no KYC, no credit card).

Available quantizations

File Quant Size Quality
Qwen3-4B-Q4_K_M.gguf Q4_K_M ~2.4 GB Best balance โ€” recommended default
Qwen3-4B-Q8_0.gguf Q8_0 ~4.7 GB Near-lossless

Quick start (llama.cpp)

# Download
huggingface-cli download QubaxAI/Qwen3-4B-GGUF Qwen3-4B-Q4_K_M.gguf --local-dir .

# Run (llama.cpp)
llama-server -m Qwen3-4B-Q4_K_M.gguf --port 8080

Or with Ollama:

ollama run hf.co/QubaxAI/Qwen3-4B-GGUF:Q4_K_M

Why Qwen3-4B?

  • Hybrid thinking / non-thinking modes in one model
  • Strong reasoning + multilingual performance for its size
  • Runs on ~4 GB VRAM at Q4 โ€” fine on a laptop

Who are we?

Qubax AI is an AI model marketplace: 340+ models, one OpenAI-compatible API, priced up to 99% below standard retail. Crypto payments, no subscription, no KYC.

License

Apache 2.0 (inherited from the base model). Model by the Qwen team โ€” this repo only redistributes quantized weights.

Downloads last month
66
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for QubaxAI/Qwen3-4B-GGUF

Finetuned
Qwen/Qwen3-4B
Quantized
(321)
this model