Instructions to use mlx-community/embeddinggemma-2-bf16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/embeddinggemma-2-bf16 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download mlx-community/embeddinggemma-2-bf16 --local-dir embeddinggemma-2-bf16
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
mlx-community/embeddinggemma-2-bf16
MLX conversion of Google EmbeddingGemma 2 in BF16 format, retaining text, image, audio, and video encoders. The model produces normalized 768-dimensional embeddings. See the original model card for intended use, training information, evaluation, and limitations.
All floating-point model weights are stored in BF16.
Weight storage: 1.489 GB (decimal). Source revision: 914f7f89142e33e77833254d9c9b90c3cef7303b. Converted with MLX-VLM revision 3d87e884, from branch pc/embeddinggemma-2, and MLX 0.32.3.
Text embeddings
Install the implementation revision that includes EmbeddingGemma 2 support:
pip install "git+https://github.com/Blaizzy/mlx-vlm.git@3d87e88402f307efbf68e568971aa887ee7d9ed0"
pip install -U "mlx>=0.32.3" "transformers>=5.18.0"
import mlx.core as mx
from transformers import AutoTokenizer
from mlx_vlm.embedding_loader import load_embedding_model
from mlx_vlm.utils import get_model_path
path = get_model_path("mlx-community/embeddinggemma-2-bf16")
model = load_embedding_model(path)
tokenizer = AutoTokenizer.from_pretrained(path)
inputs = tokenizer([
"task: search result | query: Which planet is known as the Red Planet?",
"title: none | text: Mars is known as the Red Planet.",
"title: none | text: Venus is often called Earth's twin.",
], padding=True, return_tensors="np")
embeddings = model(**{key: mx.array(value) for key, value in inputs.items()}).text_embeds
print(embeddings[:1] @ embeddings[1:].T)
# Optional Matryoshka truncation: use 128, 256, 512, or 768 dimensions.
embeddings = embeddings[:, :256]
embeddings /= mx.linalg.norm(embeddings, axis=-1, keepdims=True)
Apply the appropriate task prefixes from config_sentence_transformers.json. Queries and documents must use the same embedding dimension. Keep non-quantized weights and activations in BF16; do not cast this model to float16.
Multimodal preprocessing
All modality weights and processor configuration files are included. Image, audio, and video preprocessing requires a Transformers build that exposes EmbeddingGemma2Processor. Validation used the model's supplied upstream Transformers 5.18.0.dev0 build; stock PyPI Transformers 5.18.0 does not yet expose this processor. The text-only example above uses AutoTokenizer and does not require that processor. With a compatible processor build:
from mlx_vlm import load
model, processor = load("mlx-community/embeddinggemma-2-bf16")
inputs = processor(images=[image], return_tensors="np") # image is a PIL image
embedding = model(**{key: mx.array(value) for key, value in inputs.items()}).text_embeds
Conversion checks
All checked outputs were finite, unit-normalized, and had 768 dimensions. The text retrieval smoke test ranked the Mars passage above Venus. Checks cover six multilingual text inputs and synthetic image, audio, two-frame video, and text+image inputs; they are numerical smoke checks, not MTEB or a retrieval-quality benchmark. Cosine and absolute errors compare against the original checkpoint running in PyTorch FP32.
| Input | Minimum cosine vs FP32 | Maximum absolute error |
|---|---|---|
| audio | 0.999947 | 0.001258 |
| image | 0.999954 | 0.001135 |
| text | 0.999937 | 0.001468 |
| text_image | 0.999905 | 0.001864 |
| video | 0.999900 | 0.001688 |
BF16 conversion and reload were bit-identical to the original MLX BF16 checkpoint on all checked inputs.
Reproduce conversion
Using the implementation and compatible processor environment described above:
python -m mlx_vlm convert --hf-path google/embeddinggemma-2 --revision 914f7f89142e33e77833254d9c9b90c3cef7303b --mlx-path embeddinggemma-2-bf16 --dtype bfloat16
The source model declares Apache-2.0 licensing. This conversion preserves that license and attribution to Google.
- Downloads last month
- -
Quantized
Model tree for mlx-community/embeddinggemma-2-bf16
Base model
google/embeddinggemma-2