Required usage: pooling + instruction

Qwen3-Embedding uses last-token pooling and is instruction-aware โ€” both are required for correct results:

  • Serve with last-token pooling: llama-server -m <model>.gguf --embedding --pooling last
  • Prepend a task instruction to queries (not documents): Instruct: {task description}\nQuery:{your query}
  • Inputs should end with the EOS token <|endoftext|> (recent llama.cpp appends this automatically).

Without last-token pooling and the instruction prefix, retrieval quality collapses. Instructions improve results ~1-5%. The MTEB score below (SciFact 0.6982 nDCG@10) was measured with both. Native dimension 1024, Matryoshka (MRL) custom dimensions supported, multilingual (100+ languages), 8192-token context.

Quant note: Q4_K_M drift dips to ~0.945 min โ€” prefer Q5_K_M or Q8_0 for best fidelity.

Qwen3-Embedding-0.6B โ€” Embedding GGUF (quantization-verified)

Quantized embedding model in GGUF, served in --embedding mode via llama.cpp. This is an encoder โ€” it outputs vectors, not text. It is validated for retrieval quality and quantization fidelity, not chat behavior.

Files

  • Qwen3-Embedding-0.6B-Q4_K_M.gguf (396.5 MB)
  • Qwen3-Embedding-0.6B-Q5_K_M.gguf (444.2 MB)
  • Qwen3-Embedding-0.6B-Q8_0.gguf (639.2 MB)

Quantization drift (vs f16)

Mean cosine similarity of embeddings vs the f16 baseline. 1.0 = identical.

Quant Mean cosine Min cosine Verdict
Q4_K_M 0.97469 0.94327 good (>0.97)
Q5_K_M 0.99117 0.97828 excellent (>0.99)
Q8_0 0.99934 0.99887 excellent (>0.99)

Per-domain fidelity at Q4_K_M (which content types the quant preserves best):

Domain Mean cosine Min
science 0.96758 0.94327
legal 0.96894 0.96257
long_form 0.97152 0.97059
code 0.97431 0.96263
everyday 0.97454 0.9666
medical 0.9754 0.97154
finance 0.97994 0.97184
short_queries 0.98367 0.98154

Retrieval sanity (lightweight)

Built-in 12-query retrieval check (no external corpus): top-1 accuracy 1.0, MRR 1.0. healthy (top-1 >= 0.9)

Retrieval (MTEB)

Standardized MTEB retrieval scores (main metric, usually nDCG@10 โ€” higher is better). These are comparable across models on the MTEB leaderboard.

Task Score
SciFact 0.6982

Metric: main_score (retrieval tasks: nDCG@10). Measured on the Q8_0 quant served via llama.cpp.

Dense-retrieval mode. These scores are for standard single-vector dense retrieval (what llama.cpp serves). Models like BGE-M3 that also support sparse/multi-vector (ColBERT) modes score higher in hybrid setups โ€” that capability isn't exercised here, so compare this number against other models' dense scores, not hybrid ones.

What this is NOT

This card carries no safety, red-team, or viewpoint scores: those do not apply to an embedding model. For chat-model governance cards, see the SmartTasks text-LLM line.

Downloads last month
172
GGUF
Model size
0.6B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for smarttasks/Qwen3-Embedding-0.6B-GGUF

Quantized
(251)
this model