L1-Qwen3-8B-Exact GGUF

GGUF quantizations of l3lab/L1-Qwen3-8B-Exact.

Quantization Variants

File Bits Size Use Case
model_f16.gguf 16 ~16 GB Reference, maximum quality
model_q8_0.gguf 8 ~8.5 GB High quality, still large
model_q6_k.gguf 6 ~6.5 GB Balanced quality/size
model_q5_k_m.gguf 5 ~5.3 GB Good quality, recommended
model_q5_k_s.gguf 5 ~5.1 GB Slightly faster Q5
model_q4_k_m.gguf 4 ~4.3 GB Standard quantization
model_q4_k_s.gguf 4 ~4.1 GB Slightly faster Q4
model_q3_k_l.gguf 3 ~3.3 GB Small model variant
model_q3_k_m.gguf 3 ~3.2 GB Balanced 3-bit
model_q3_k_s.gguf 3 ~3.0 GB Fast 3-bit
model_q2_k.gguf 2 ~2.3 GB Minimal size, CPU-friendly

VRAM Requirements

Quantization GPU VRAM RAM (CPU)
F16 ~16 GB ~20 GB
Q8_0 ~9 GB ~11 GB
Q6_K ~7 GB ~9 GB
Q5_K_M ~6 GB ~8 GB
Q4_K_M ~5 GB ~6 GB
Q3_K_M ~3.5 GB ~5 GB
Q2_K ~2.5 GB ~3.5 GB

Usage

llama.cpp

./main -m model_q5_k_m.gguf -n 256 -p "Your prompt here"

llama-cpp-python

from llama_cpp import Llama

llm = Llama(
    model_path="model_q5_k_m.gguf",
    n_gpu_layers=-1,  # Use GPU
    n_ctx=2048,
)

response = llm("Your prompt here", max_tokens=256)
print(response["choices"][0]["text"])

LM Studio

  1. Download any .gguf file from this repo
  2. Open in LM Studio
  3. Configure context and parameters
  4. Start chat

Ollama

ollama create l1-qwen3-8b -f Modelfile

Modelfile:

FROM model_q5_k_m.gguf
PARAMETER temperature 0.7

Base Model

l3lab/L1-Qwen3-8B-Exact โ€” Fine-tuned reasoning model based on Qwen3 8B, trained with reinforcement learning to optimize reasoning step count.

License

Apache 2.0 (as per base model)

Downloads last month
1,369
GGUF
Model size
8B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for tinyopsec/L1-Qwen3-8B-Exact-GGUF

Quantized
(2)
this model