EmbeddingGemma-2 (4-bit weights)
Model Overview
EmbeddingGemma-2 is a multimodal embedding model from Google. It maps text, images, video and audio into one 768-dimensional embedding space (Matryoshka: 768, 512, 256 or 128 dimensions), so a text query can be compared directly with photos, video frames and sound clips.
This build has been optimized for the Synaptics Astra™ SL2610-Series processors with Torq NPU:
4-bit block-quantized weights (block 32, from the onnx-community *_q4 export) with bf16 activations. The bf16 build (closest to the original model, more memory) is on the bf16 branch.
Model Features
- Model Type: Multimodal embedding model
- Input: Text, image, video, audio
- Output: 768-d embedding
Files
| File | Runs on | Purpose |
|---|---|---|
text_body_s128.vmfb, text_body_s512.vmfb |
NPU | Text body for 128 / 512 tokens (queries, images; documents, video, audio) |
vision_384x384.vmfb |
NPU | Vision encoder: a 384x384 image or video frame -> 64 soft tokens |
audio_140.vmfb … audio_1120.vmfb |
NPU | Audio encoder for 1.4 / 2.8 / 5.6 / 11.2 s windows |
token_embeddings.npy |
CPU | Token-embedding table, memory-mapped |
mel_filters.npy, mel_window.npy |
CPU | Audio log-mel front end |
tokenizer.json, config.json, embeddinggemma2_manifest.json |
CPU | Tokenizer, model config, static shapes and NPU memory per network |
onnx/ |
— | Source ONNX graphs (fp32 and 4-bit) that torq-tools builds these files from: a copy of onnx-community/embeddinggemma-2-ONNX (on main) |
Performance (SL2619, one network per process)
| Network | Latency | RAM | File |
|---|---|---|---|
| text S=128 | 549 ms | 122 MB | 87 MB |
| text S=512 | 2825 ms | 140 MB | 103 MB |
| vision 384x384 | 3310 ms | 157 MB | 122 MB |
| audio 1.4 s | 279 ms | 210 MB | 180 MB |
| audio 2.8 s | 600 ms | 211 MB | 181 MB |
| audio 5.6 s | 1337 ms | 217 MB | 186 MB |
| audio 11.2 s | 2670 ms | 227 MB | 196 MB |
ONNX Runtime with the same 4-bit weights on the board CPU (2 threads) is 5.9–11.3x slower (text S=128 3.7 s, vision 21.6 s, audio 11.2 s 23.3 s).
Deployment
Usage and demos (text/image/video/audio similarity, and search over a media library) are in
torq-examples under
Google/Google-EmbeddingGemma-2:
python setup_demos.py Google-EmbeddingGemma-2
Learn More
- Synaptics AI Developer Zone: Get started with documentation, tutorials and resources for your Edge AI journey.
- Astra Support Portal: Connect with our engineering team and community.
- Downloads last month
- 19
Model tree for Synaptics/Google-EmbeddingGemma-2
Base model
google/embeddinggemma-2