qwen3-embedding-0.6b-int8-ov

OpenVINO IR export of Qwen/Qwen3-Embedding-0.6B, quantized to INT8 — 590 MB.

Quantization weight compression (optimum)
Lesson 11 Embedding search / 12 RAG
Built by convert/convert_all.py of the ARCademy OpenVINO courseware

Last-token pooling over last_hidden_state, then L2-normalise. Queries take an 'Instruct: …\nQuery: …' prefix.

from huggingface_hub import snapshot_download
model_dir = snapshot_download("circulus/qwen3-embedding-0.6b-int8-ov")
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for circulus/qwen3-embedding-0.6b-int8-ov

Finetuned
(276)
this model