Qwen3-0.6B mEinstein GGUF

Fine-tuned version of Qwen/Qwen3-0.6B on the mEinstein personal AI assistant dataset.

Available Quantizations

File Format Size Use case
Qwen_mEinstein_F16.gguf F16 ~1.2 GB Highest quality
Qwen_mEinstein_Q8_0.gguf Q8_0 ~0.6 GB Balanced
Qwen_mEinstein_Q4_K_M.gguf Q4_K_M ~0.4 GB Fastest / lightest

How to Run

./llama-cli -m Qwen_mEinstein_Q4_K_M.gguf -p "Who am I?" -n 256

Training Details

  • Base model: Qwen/Qwen3-0.6B
  • Method: QLoRA (4-bit) + SFT
  • LoRA target modules: q_proj, k_proj, v_proj, o_proj
  • Dataset: mEinstein personal assistant dataset
Downloads last month
30
GGUF
Model size
0.6B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for nikhil-singh-iphtech/qwen_v3_gguf

Finetuned
Qwen/Qwen3-0.6B
Quantized
(379)
this model