qwen3-embedding-0.6b-int8-ov
OpenVINO IR export of Qwen/Qwen3-Embedding-0.6B, quantized to INT8 — 590 MB.
| Quantization | weight compression (optimum) |
| Lesson | 11 Embedding search / 12 RAG |
| Built by | convert/convert_all.py of the ARCademy OpenVINO courseware |
Last-token pooling over last_hidden_state, then L2-normalise. Queries take an 'Instruct: …\nQuery: …' prefix.
from huggingface_hub import snapshot_download
model_dir = snapshot_download("circulus/qwen3-embedding-0.6b-int8-ov")
- Downloads last month
- -
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support