EmbeddingGemma-2 (4-bit weights)

Model Overview

EmbeddingGemma-2 is a multimodal embedding model from Google. It maps text, images, video and audio into one 768-dimensional embedding space (Matryoshka: 768, 512, 256 or 128 dimensions), so a text query can be compared directly with photos, video frames and sound clips.

This build has been optimized for the Synaptics Astra™ SL2610-Series processors with Torq NPU: 4-bit block-quantized weights (block 32, from the onnx-community *_q4 export) with bf16 activations. The bf16 build (closest to the original model, more memory) is on the bf16 branch.

Model Features

  • Model Type: Multimodal embedding model
  • Input: Text, image, video, audio
  • Output: 768-d embedding

Files

File Runs on Purpose
text_body_s128.vmfb, text_body_s512.vmfb NPU Text body for 128 / 512 tokens (queries, images; documents, video, audio)
vision_384x384.vmfb NPU Vision encoder: a 384x384 image or video frame -> 64 soft tokens
audio_140.vmfb … audio_1120.vmfb NPU Audio encoder for 1.4 / 2.8 / 5.6 / 11.2 s windows
token_embeddings.npy CPU Token-embedding table, memory-mapped
mel_filters.npy, mel_window.npy CPU Audio log-mel front end
tokenizer.json, config.json, embeddinggemma2_manifest.json CPU Tokenizer, model config, static shapes and NPU memory per network
onnx/ — Source ONNX graphs (fp32 and 4-bit) that torq-tools builds these files from: a copy of onnx-community/embeddinggemma-2-ONNX (on main)

Performance (SL2619, one network per process)

Network Latency RAM File
text S=128 549 ms 122 MB 87 MB
text S=512 2825 ms 140 MB 103 MB
vision 384x384 3310 ms 157 MB 122 MB
audio 1.4 s 279 ms 210 MB 180 MB
audio 2.8 s 600 ms 211 MB 181 MB
audio 5.6 s 1337 ms 217 MB 186 MB
audio 11.2 s 2670 ms 227 MB 196 MB

ONNX Runtime with the same 4-bit weights on the board CPU (2 threads) is 5.9–11.3x slower (text S=128 3.7 s, vision 21.6 s, audio 11.2 s 23.3 s).

Deployment

Usage and demos (text/image/video/audio similarity, and search over a media library) are in torq-examples under Google/Google-EmbeddingGemma-2:

python setup_demos.py Google-EmbeddingGemma-2

Learn More

Downloads last month
19
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Synaptics/Google-EmbeddingGemma-2

Quantized
(59)
this model