Qwen3-1.7B GGUF โ€” run it anywhere

Ready-to-run GGUF build of Qwen/Qwen3-1.7B for llama.cpp, Ollama, and LM Studio. At ~1.8 GB, it runs on CPUs and laptops โ€” no GPU required.

Need frontier-scale models instead? Our API at qubax.ai serves 340+ models through one OpenAI-compatible endpoint at some of the lowest prices on the market (pay with crypto, no KYC, no credit card).

Files

File Quant Size
Qwen3-1.7B-Q8_0.gguf Q8_0 ~1.8 GB

Quick start

# llama.cpp
llama-server -m Qwen3-1.7B-Q8_0.gguf --port 8080
# Ollama
ollama run hf.co/QubaxAI/Qwen3-1.7B-GGUF:Q8_0

About

Qubax AI is an AI model marketplace: 340+ models, one OpenAI-compatible API, priced up to 99% below standard retail. Crypto payments, no subscription, no KYC.

Downloads last month
33
GGUF
Model size
2B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for QubaxAI/Qwen3-1.7B-GGUF

Finetuned
Qwen/Qwen3-1.7B
Quantized
(380)
this model