qwen3-0.6b-int4-ov

OpenVINO IR export of Qwen/Qwen3-0.6B, quantized to INT4 โ€” 387 MB.

Quantization weight compression, group 128, symmetric
Lesson 09 Speculative decoding (draft)
Built by convert/convert_all.py of the ARCademy OpenVINO courseware

Use as openvino_genai.draft_model(model_dir, device) beside qwen3-4b-int4-ov โ€” same tokenizer family, which speculative decoding requires.

from huggingface_hub import snapshot_download
model_dir = snapshot_download("circulus/qwen3-0.6b-int4-ov")
Downloads last month
11
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for circulus/qwen3-0.6b-int4-ov

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1257)
this model