Instructions to use ngquocvinh/EmbeddingGemma-2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use ngquocvinh/EmbeddingGemma-2-GGUF with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("ngquocvinh/EmbeddingGemma-2-GGUF") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ngquocvinh/EmbeddingGemma-2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ngquocvinh/EmbeddingGemma-2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ngquocvinh/EmbeddingGemma-2-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ngquocvinh/EmbeddingGemma-2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ngquocvinh/EmbeddingGemma-2-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ngquocvinh/EmbeddingGemma-2-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ngquocvinh/EmbeddingGemma-2-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ngquocvinh/EmbeddingGemma-2-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ngquocvinh/EmbeddingGemma-2-GGUF:Q4_K_M
Use Docker
docker model run hf.co/ngquocvinh/EmbeddingGemma-2-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use ngquocvinh/EmbeddingGemma-2-GGUF with Ollama:
ollama run hf.co/ngquocvinh/EmbeddingGemma-2-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use ngquocvinh/EmbeddingGemma-2-GGUF with Docker Model Runner:
docker model run hf.co/ngquocvinh/EmbeddingGemma-2-GGUF:Q4_K_M
- Lemonade
How to use ngquocvinh/EmbeddingGemma-2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ngquocvinh/EmbeddingGemma-2-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.EmbeddingGemma-2-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
EmbeddingGemma 2 GGUF
Community GGUF quantizations of google/embeddinggemma-2.
Send a coffee โ
I build and test these releases myself. Your coffee helps keep me going.
Thank you for supporting this work.
About EmbeddingGemma 2
EmbeddingGemma 2 is a multilingual, multimodal embedding model from Google DeepMind. It maps text and code, images, video, and audio into a shared 768-dimensional vector space and has an 8,192-token shared context window. The upstream checkpoint has a 270M-parameter text path plus optional vision and audio encoders. This release contains a quantized text backbone and a shared BF16 vision/audio projector. Text, image, and audio inputs returned normalized 768-dimensional vectors in local llama-server CPU smoke checks. See the official model card for supported task prefixes, modalities, and input guidance.
Embedding evaluation
Every value below comes from local measurements of the locked BF16 GGUF and quantized files; the numbers are not copied from the upstream model card. The fixed task evaluation uses the 1,379-pair test split of MTEB STSBenchmark STS, dataset revision 96943a16ea6a35129e253c659081cb59daf81b30. It covers 2,552 unique sentences with the task: sentence similarity | query: prefix, an 8,192-token context, and the CPU llama.cpp runtime at commit 9c2e0e491a822adae1f0b1c831adb4160057d24f. Spearman and Pearson measure correlation between cosine similarity and human scores. Mean and 5th-percentile cosine measure vector agreement with the BF16 GGUF reference. Higher values indicate stronger task correlation or closer vector agreement; smaller files use less disk. This is a task-specific embedding evaluation, not the full MTEB benchmark. Token-level KLD, next-token Top-1, PPL, and RMS probability deltas do not apply to this embedding model.
There are 26 quantized text artifacts in this release. Hold-out measurements are available for 24; the two TQ variants passed runtime smoke checks but were not included in this evaluation checkpoint, so no quality results are claimed for them.
BF16 reference baseline
| Reference | Size (GB) | Spearman vs human โ | Pearson vs human โ | Mean cosine vs BF16 | 5th-percentile cosine vs BF16 |
|---|---|---|---|---|---|
| Local BF16 text GGUF | 0.557950 | 0.875414 | 0.865994 | 1.000000 | 1.000000 |
Quantized variants compared with BF16
| File | Size (GB) | Spearman vs human โ | Pearson vs human โ | Mean cosine vs BF16 โ | 5th-percentile cosine vs BF16 โ |
|---|---|---|---|---|---|
EmbeddingGemma-2-Q8_0.gguf |
0.309856 | 0.875125 | 0.865776 | 0.999930 | 0.999903 |
EmbeddingGemma-2-Q6_K.gguf |
0.245765 | 0.874311 | 0.864975 | 0.999548 | 0.999373 |
EmbeddingGemma-2-Q5_K_M.gguf |
0.212759 | 0.874081 | 0.864863 | 0.999190 | 0.998841 |
EmbeddingGemma-2-Q5_K_S.gguf |
0.210670 | 0.874892 | 0.865614 | 0.999104 | 0.998727 |
EmbeddingGemma-2-Q4_K_M.gguf |
0.181695 | 0.872686 | 0.864659 | 0.997824 | 0.996776 |
EmbeddingGemma-2-Q4_K_S.gguf |
0.178164 | 0.873112 | 0.864908 | 0.997379 | 0.996201 |
EmbeddingGemma-2-IQ4_NL.gguf |
0.177640 | 0.871242 | 0.862865 | 0.996884 | 0.995270 |
EmbeddingGemma-2-IQ4_XS.gguf |
0.169383 | 0.871200 | 0.863202 | 0.996687 | 0.994900 |
EmbeddingGemma-2-Q3_K_L.gguf |
0.154440 | 0.869874 | 0.860846 | 0.993184 | 0.989422 |
EmbeddingGemma-2-Q3_K_M.gguf |
0.148870 | 0.869604 | 0.860705 | 0.992364 | 0.988323 |
EmbeddingGemma-2-IQ3_M.gguf |
0.145749 | 0.870880 | 0.862364 | 0.992638 | 0.988771 |
EmbeddingGemma-2-Q3_K_S.gguf |
0.142546 | 0.866108 | 0.853203 | 0.987335 | 0.981859 |
EmbeddingGemma-2-IQ3_S.gguf |
0.142546 | 0.870980 | 0.863233 | 0.990906 | 0.986058 |
EmbeddingGemma-2-IQ3_XS.gguf |
0.139793 | 0.867124 | 0.859804 | 0.989128 | 0.983518 |
EmbeddingGemma-2-IQ3_XXS.gguf |
0.135776 | 0.862974 | 0.858534 | 0.983963 | 0.976505 |
EmbeddingGemma-2-IQ2_M.gguf |
0.130910 | 0.862532 | 0.861373 | 0.971570 | 0.960803 |
EmbeddingGemma-2-IQ2_S.gguf |
0.127600 | 0.856632 | 0.851405 | 0.963077 | 0.949864 |
EmbeddingGemma-2-Q2_K.gguf |
0.120394 | 0.856144 | 0.854341 | 0.972891 | 0.960440 |
EmbeddingGemma-2-Q2_K_S.gguf |
0.116446 | 0.843346 | 0.843931 | 0.962411 | 0.948042 |
EmbeddingGemma-2-IQ2_XS.gguf |
0.110946 | 0.846829 | 0.838447 | 0.944608 | 0.927359 |
EmbeddingGemma-2-IQ2_XXS.gguf |
0.107178 | 0.834294 | 0.827123 | 0.919032 | 0.899266 |
EmbeddingGemma-2-IQ1_M.gguf |
0.103041 | 0.797642 | 0.801111 | 0.819494 | 0.778721 |
EmbeddingGemma-2-IQ1_S.gguf |
0.100558 | 0.761951 | 0.772373 | 0.768243 | 0.728349 |
EmbeddingGemma-2-Q1_0.gguf |
0.066163 | 0.061457 | 0.062574 | 0.513763 | 0.425907 |
Q5_K_S and Q4_K_S are the current memory/fidelity candidates on this measured task: their Spearman scores are 0.874892 and 0.873112 with sizes 0.210670 GB and 0.178164 GB. No VRAM-fit or throughput claim is made. Q1_0 and IQ1_S return valid vectors but show substantial loss on this task; the TQ formats have smoke results only. The machine-readable evaluation summary and reproducibility manifest include the measurements and method. The Q8_0 artifact and BF16 vision/audio projector also passed local text, image, and audio embedding smoke checks.
Quick start
Start an embedding endpoint with llama.cpp:
llama-server \
-m EmbeddingGemma-2-Q8_0.gguf \
--mmproj mmproj-EmbeddingGemma-2-BF16.gguf \
--embedding \
--pooling mean \
--ctx-size 8192 \
--port 8080
Send an English search query using the model's task prefix:
curl http://localhost:8080/v1/embeddings \
-H 'Content-Type: application/json' \
-d '{"model":"EmbeddingGemma-2","input":"task: search result | query: What causes the northern lights?"}'
For text-only workloads, omit --mmproj. For documents, use title: {title} | text: {content} or title: none | text: {content} when there is no title. Keep query and document vectors at the same dimension. The server returns normalized vectors by default.
Reproducibility and validation
See the reproducibility manifest, artifact measurements, and SHA256 checksums. The manifest records the locked upstream revision and model hash, BF16 GGUF source, converter/runtime revision, model-specific calibration/imatrix, quantization commands, and smoke profile. Raw conversion, imatrix, quantization, and validation logs are kept locally and are not part of this repository.
License and attribution
The upstream model is licensed under Apache 2.0. These are community GGUF quantizations, not an official Google or Google DeepMind release or endorsement.
- Downloads last month
- 4,365
1-bit
2-bit
3-bit
4-bit
5-bit
6-bit
8-bit
Model tree for ngquocvinh/EmbeddingGemma-2-GGUF
Base model
google/embeddinggemma-2