Qwen3-1.7B-Coder-Distilled-SFT — GGUF

GGUF quantizations of reaperdoesntknow/Qwen3-1.7B-Coder-Distilled-SFT, quantized by tinyopsec.

A 1.7B model built in two stages: knowledge distillation from Qwen3-Coder-30B-A3B-Instruct (30B MoE teacher) to establish a structured STEM reasoning backbone, then SFT on ~54,600 logical inference problems. Architecture: Qwen3ForCausalLM.


Available Quantizations

File Bits Approx. Size Use Case
model_f16.gguf 16 ~3.4 GB Full precision, reference
model_q8_0.gguf 8 ~1.8 GB Best quality, fits in RAM easily
model_q6_k.gguf 6 ~1.4 GB Excellent quality, recommended
model_q5_k_m.gguf 5 ~1.2 GB Great balance quality/size
model_q5_k_s.gguf 5 ~1.1 GB Slightly smaller Q5 variant
model_q4_k_m.gguf 4 ~1.0 GB Good quality, very portable
model_q4_k_s.gguf 4 ~0.95 GB Smaller Q4 variant
model_q3_k_l.gguf 3 ~0.85 GB Lower quality, minimal RAM
model_q3_k_m.gguf 3 ~0.80 GB Minimal footprint
model_q3_k_s.gguf 3 ~0.75 GB Smallest Q3 variant
model_q2_k.gguf 2 ~0.60 GB Extreme compression, lowest quality

VRAM Requirements

Quantization VRAM
F16 ~3.4 GB
Q8_0 ~1.8 GB
Q6_K ~1.4 GB
Q5_K_M ~1.2 GB
Q4_K_M ~1.0 GB
Q3_K_M ~0.80 GB
Q2_K ~0.60 GB

Usage

llama.cpp

./llama-cli -m model_q4_k_m.gguf -p "### Instruction:\nWhat can we infer from: All cats are mammals. Whiskers is a cat.\n\n### Response:" -n 256

llama-cpp-python

from llama_cpp import Llama

llm = Llama(model_path="model_q4_k_m.gguf", n_ctx=1024)
output = llm(
    "### Instruction:\nWhat can we infer from: All cats are mammals. Whiskers is a cat.\n\n### Response:",
    max_tokens=256,
    stop=["### Instruction:"],
)
print(output["choices"][0]["text"])

LM Studio

Download any .gguf file from this repo and load it directly in LM Studio.

Ollama

ollama run hf.co/tinyopsec/Qwen3-1.7B-Coder-Distilled-SFT-GGUF

Prompt Formats

Logical inference (Stage 2 — primary):

### Instruction:
[Your question or logical inference problem]

### Response:

STEM derivation (Stage 1 — also supported):

Solve the following problem carefully and show a rigorous derivation.

Problem:
[Your problem]

Proof:

Model Details

Attribute Value
Architecture Qwen3ForCausalLM
Parameters ~2B (1.7B effective)
Base model Qwen/Qwen3-1.7B
Teacher model Qwen/Qwen3-Coder-30B-A3B-Instruct
Context length 1024 tokens (training)
Precision BF16
License Apache 2.0

Good for: Logical inference, propositional logic, formal reasoning, STEM derivation, structured argumentation, educational tutoring, edge deployment.

Not for: General code generation, formal proof verification (use Lean/Coq), safety-critical tasks, or long context beyond 1024 tokens.


Original Model

reaperdoesntknow/Qwen3-1.7B-Coder-Distilled-SFT by Convergent Intelligence LLC: Research Division.

Downloads last month
1,368
GGUF
Model size
2B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tinyopsec/Qwen3-1.7B-Coder-Distilled-SFT-GGUF

Finetuned
Qwen/Qwen3-1.7B
Quantized
(2)
this model