Qwen3.8-4B-Empero-AI-FullStack — GGUF Quantizations

This repository contains GGUF quantizations of iBotIA/Qwen3.8-4B-Empero-AI-FullStack.


📦 Available Quantizations

File Bits Size (approx.) Use case
model_f16.gguf 16-bit ~8.7 GB Maximum quality, reference
model_q8_0.gguf 8-bit ~4.7 GB Near-lossless, high VRAM
model_q6_k.gguf 6-bit ~3.6 GB Excellent quality
model_q5_k_m.gguf 5-bit ~3.1 GB Great quality/size balance
model_q5_k_s.gguf 5-bit ~3.0 GB Slightly smaller than K_M
model_q4_k_m.gguf 4-bit ~2.5 GB Recommended default
model_q4_k_s.gguf 4-bit ~2.4 GB Smaller 4-bit variant
model_q3_k_l.gguf 3-bit ~2.1 GB Low VRAM, decent quality
model_q3_k_m.gguf 3-bit ~1.9 GB Balanced 3-bit
model_q3_k_s.gguf 3-bit ~1.7 GB Minimum 3-bit
model_q2_k.gguf 2-bit ~1.3 GB Extreme compression

IQ quants (IQ4_XS, etc.) coming soon — require imatrix calibration.


🚀 Usage

llama.cpp

./llama-cli -m model_q4_k_m.gguf -p "Your prompt here" -n 512

LM Studio

Download any .gguf file and load it directly in LM Studio.

Ollama

ollama run hf.co/tinyopsec/Qwen3.8-4B-Empero-AI-FullStack-GGUF:Q4_K_M

Python (llama-cpp-python)

from llama_cpp import Llama

llm = Llama.from_pretrained(
    repo_id="tinyopsec/Qwen3.8-4B-Empero-AI-FullStack-GGUF",
    filename="model_q4_k_m.gguf",
)
output = llm("Your prompt here", max_tokens=512)
print(output["choices"][0]["text"])

🔧 Quantization Details


💡 Which quant should I use?

VRAM Recommended
2 GB Q2_K
3 GB Q3_K_M
4 GB Q4_K_M ✅
6 GB Q5_K_M
8 GB Q6_K
12 GB+ Q8_0 / F16

📄 License

Refer to the original model license.

Downloads last month
1,817
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tinyopsec/Qwen3.8-4B-Empero-AI-FullStack-GGUF

Finetuned
Qwen/Qwen3.5-4B
Quantized
(3)
this model

Collection including tinyopsec/Qwen3.8-4B-Empero-AI-FullStack-GGUF