Instructions to use litert-community/embeddinggemma-2-text-270m-litert-lm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT-LM
How to use litert-community/embeddinggemma-2-text-270m-litert-lm with LiteRT-LM:
# LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM) # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter). # For platform-specific integration guides, please refer to the official developer website: # https://ai.google.dev/edge/litert-lm # To try LiteRT-LM, the easiest way is to use our CLI tool. # 1. Install the LiteRT-LM CLI tool: pip install -U litert-lm # 2. Download and run this model locally: # See: https://ai.google.dev/edge/litert-lm/cli # A single .litertlm file in the repo is picked automatically; otherwise the CLI asks which one to run # (or pass its name right after the repo id). litert-lm run \ --from-huggingface-repo=litert-community/embeddinggemma-2-text-270m-litert-lm \ --prompt="Write me a poem"
- Notebooks
- Google Colab
- Kaggle
litert-community/embeddinggemma-2-text-270m-litert-lm
Main Model Card: google/embeddinggemma-2
EmbeddingGemma 2 Text 270M is a highly efficient, text-only variant of the EmbeddingGemma 2 model family, optimized for low-latency deployment on consumer hardware and edge devices via LiteRT. With a footprint of 270M parameters (combining a 130M transformer with a 140M embedder), this model delivers high-quality text representations. It supports over 100 languages, programming code, and features Matryoshka Representation Learning (MRL) for adjustable vector dimensions (down to 128d) to dramatically lower vector storage overhead. This variant is ideal for text-centric on-device applications, including document search, classification, clustering, and question answering.
Try EmbeddingGemma 2 with LiteRT-LM
Ready to integrate this into your product? Get started here.
Try EmbeddingGemma 2 with MediaPipe
MediaPipe uses EmbeddingGemma 2 and LiteRT-LM to enable high-level, cross-platform tasks which you can experience firsthand with the following interactive web demos:
- Semantic image search powered by MediaPipe Universal Embedder and Semantic Retriever task.
- Decision making powered by MediaPipe Decision Maker.
Model Specifications
| EmbeddingGemma 2 Text 270M |
EmbeddingGemma 2 Text Vision 440M |
EmbeddingGemma 2 740M |
|
|---|---|---|---|
| Supported Modalities | Text | Text, Images | Text, Images, Video, Audio |
| Parameters | Total: 270M Transformer: 130M Embeddings: 140M |
Total: 440M Transformer: 130M Embeddings: 140M Vision Encoder: 170M |
Total: 740M Transformer: 130M Embeddings: 140M Vision Encoder: 170M Audio Encoder: 300M |
| Quantization Scheme | Transformer: int4 per-channel (QAT) Embeddings: int4 per-channel (QAT) |
Transformer: int4 per-channel (QAT) Embeddings: int4 per-channel (QAT) Vision Encoder: int8 per-channel (QAT) |
Transformer: int4 per-channel (QAT) Embeddings: int4 per-channel (QAT) Vision Encoder: int8 per-channel (QAT) Audio Encoder: mixed int2/int4/int8 per-channel (QAT) |
| Download Size (CPU/GPU file) | 165 MB | 388 MB | 485 MB |
| On-Demand Modality Loading | - | Yes | Yes |
| Supported Input Sizes | Text tokens: 128, 256, 512, 1024, 2048, 8192 |
Text tokens: 128, 256, 512, 1024, 2048, 8192 Image soft tokens: 70, 140 |
Text tokens: 128, 256, 512, 1024, 2048, 8192 Image soft tokens: 70, 140 Audio soft tokens: 12 |
| Supported Output Sizes | 128 (512 bytes) 256 (1024 bytes) 512 (2048 bytes) 768 (3072 bytes) |
128 (512 bytes) 256 (1024 bytes) 512 (2048 bytes) 768 (3072 bytes) |
128 (512 bytes) 256 (1024 bytes) 512 (2048 bytes) 768 (3072 bytes) |
| Link | litert-community/embeddinggemma-2-text-270m-litert-lm | litert-community/embeddinggemma-2-text-vision-440m-litert-lm | litert-community/embeddinggemma-2-740m-litert-lm |
Additional Notes:
- LiteRT-LM automatically scales and patchifies images of any size to fit EmbeddingGemma 2's vision encoder model.
- LiteRT-LM supports EmbeddingGemma 2's streamed audio allowing a variety of different audio lengths.
EmbeddingGemma 2 Performance on LiteRT-LM
The text performance was measured by loading and running the 128 text signature. The latency reported is the average of 5 iterations.
Memory was measured with each platform's native metrics and is not directly comparable across operating systems. CPU memory was measured using, rusage::ru_maxrss on Android, Linux and IOT, task_vm_info::phys_footprint on iOS and MacBook, and process_memory_counters::PrivateUsage on Windows. With the exception of task_vm_info::phys_footprint, accelerator (GPU/NPU/TPU) memory is not included. Memory measurements were taken from the second load which reads from loading caches. Memory usage on the first load may vary.
Android
| Device | Backend | Text latency |
|---|---|---|
| Google Pixel 11 Pro | TPU | 8.3 ms |
| S26 Ultra | CPU | 27.1 ms |
| S26 Ultra | GPU | 25.9 ms |
| Device | Backend | Text CPU memory |
|---|---|---|
| Google Pixel 11 Pro | TPU | 112 MB |
| S26 Ultra | CPU | 334 MB |
| S26 Ultra | GPU | 333 MB |
iOS
| Device | Backend | Text latency |
|---|---|---|
| iPhone 18 Pro | CPU | 41.8 ms |
| iPhone 18 Pro | GPU | 11.6 ms |
| Device | Backend | Text CPU/GPU memory |
|---|---|---|
| iPhone 18 Pro | CPU | 84 MB |
| iPhone 18 Pro | GPU | 85 MB |
Linux
| Device | Backend | Text latency |
|---|---|---|
| Arm 2.3 & 2.8GHz | CPU | 105 ms |
| NVIDIA GeForce RTX 4090 | GPU | 7.6 ms |
| Device | Backend | Text CPU memory |
|---|---|---|
| Arm 2.3 & 2.8GHz | CPU | 310 MB |
| NVIDIA GeForce RTX 4090 | GPU | 528 MB |
macOS
| Device | Backend | Text latency |
|---|---|---|
| MacBook Pro M5 | CPU | 31.3 ms |
| MacBook Pro M5 | GPU | 9.5 ms |
| Device | Backend | Text CPU/GPU memory |
|---|---|---|
| MacBook Pro M5 | CPU | 165 MB |
| MacBook Pro M5 | GPU | 233 MB |
Windows
| Device | Backend | Text latency |
|---|---|---|
| Dell XPS 16 (Intel Core Ultra Series 3) | CPU | 71.7 ms |
| Dell XPS 16 (Intel Core Ultra Series 3) | GPU | 19.2 ms |
| Dell XPS 16 (Intel Core Ultra Series 3) | Running Intel OpenVINO NPU | 13.3 ms |
| Device | Backend | Text CPU memory |
|---|---|---|
| Dell XPS 16 (Intel Core Ultra Series 3) | CPU | 211 MB |
| Dell XPS 16 (Intel Core Ultra Series 3) | GPU | 822 MB |
| Dell XPS 16 (Intel Core Ultra Series 3) | Running Intel OpenVINO NPU | 201 MB |
Web
| Device | Backend | Text latency |
|---|---|---|
| MacBook Pro M5 | GPU | 21.8 ms |
IoT
| Device | Backend | Text latency |
|---|---|---|
| Raspberry Pi 5 16GB | CPU | 161 ms |
| Jetson Orin Nano | CPU | 207 ms |
| Jetson Orin Nano | GPU | 88 ms |
| Arduino VENTUNO Q | NPU | 13.6 ms |
| Device | Backend | Text CPU memory |
|---|---|---|
| Raspberry Pi 5 16GB | CPU | 282 MB |
| Jetson Orin Nano | CPU | 302 MB |
| Jetson Orin Nano | GPU | 523 MB |
| Arduino VENTUNO Q | NPU | 206 MB |
- Downloads last month
- 3,155
Model tree for litert-community/embeddinggemma-2-text-270m-litert-lm
Base model
google/embeddinggemma-2