qwen3-0.6b-int4-ov
OpenVINO IR export of Qwen/Qwen3-0.6B, quantized to INT4 โ 387 MB.
| Quantization | weight compression, group 128, symmetric |
| Lesson | 09 Speculative decoding (draft) |
| Built by | convert/convert_all.py of the ARCademy OpenVINO courseware |
Use as openvino_genai.draft_model(model_dir, device) beside qwen3-4b-int4-ov โ same tokenizer family, which speculative decoding requires.
from huggingface_hub import snapshot_download
model_dir = snapshot_download("circulus/qwen3-0.6b-int4-ov")
- Downloads last month
- 11
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support