LiteRT-LM

litert-community/embeddinggemma-2-text-270m-litert-lm

Main Model Card: google/embeddinggemma-2

EmbeddingGemma 2 Text 270M is a highly efficient, text-only variant of the EmbeddingGemma 2 model family, optimized for low-latency deployment on consumer hardware and edge devices via LiteRT. With a footprint of 270M parameters (combining a 130M transformer with a 140M embedder), this model delivers high-quality text representations. It supports over 100 languages, programming code, and features Matryoshka Representation Learning (MRL) for adjustable vector dimensions (down to 128d) to dramatically lower vector storage overhead. This variant is ideal for text-centric on-device applications, including document search, classification, clustering, and question answering.

Try EmbeddingGemma 2 with LiteRT-LM

Ready to integrate this into your product? Get started here.

Try EmbeddingGemma 2 with MediaPipe

MediaPipe uses EmbeddingGemma 2 and LiteRT-LM to enable high-level, cross-platform tasks which you can experience firsthand with the following interactive web demos:

Model Specifications

EmbeddingGemma 2
Text 270M
EmbeddingGemma 2
Text Vision 440M
EmbeddingGemma 2
740M
Supported Modalities Text Text, Images Text, Images, Video, Audio
Parameters Total: 270M
Transformer: 130M
Embeddings: 140M
Total: 440M
Transformer: 130M
Embeddings: 140M
Vision Encoder: 170M
Total: 740M
Transformer: 130M
Embeddings: 140M
Vision Encoder: 170M
Audio Encoder: 300M
Quantization Scheme Transformer:
    int4 per-channel (QAT)
Embeddings:
    int4 per-channel (QAT)
Transformer:
    int4 per-channel (QAT)
Embeddings:
    int4 per-channel (QAT)
Vision Encoder:
    int8 per-channel (QAT)
Transformer:
    int4 per-channel (QAT)
Embeddings:
    int4 per-channel (QAT)
Vision Encoder:
    int8 per-channel (QAT)
Audio Encoder:
    mixed int2/int4/int8
    per-channel (QAT)
Download Size (CPU/GPU file) 165 MB 388 MB 485 MB
On-Demand Modality Loading - Yes Yes
Supported Input Sizes Text tokens: 128, 256, 512,
    1024, 2048, 8192
Text tokens: 128, 256, 512,
    1024, 2048, 8192
Image soft tokens: 70, 140
Text tokens: 128, 256, 512,
    1024, 2048, 8192
Image soft tokens: 70, 140
Audio soft tokens: 12
Supported Output Sizes 128 (512 bytes)
256 (1024 bytes)
512 (2048 bytes)
768 (3072 bytes)
128 (512 bytes)
256 (1024 bytes)
512 (2048 bytes)
768 (3072 bytes)
128 (512 bytes)
256 (1024 bytes)
512 (2048 bytes)
768 (3072 bytes)
Link litert-community/embeddinggemma-2-text-270m-litert-lm litert-community/embeddinggemma-2-text-vision-440m-litert-lm litert-community/embeddinggemma-2-740m-litert-lm

Additional Notes:

  • LiteRT-LM automatically scales and patchifies images of any size to fit EmbeddingGemma 2's vision encoder model.
  • LiteRT-LM supports EmbeddingGemma 2's streamed audio allowing a variety of different audio lengths.

EmbeddingGemma 2 Performance on LiteRT-LM

The text performance was measured by loading and running the 128 text signature. The latency reported is the average of 5 iterations.

Memory was measured with each platform's native metrics and is not directly comparable across operating systems. CPU memory was measured using, rusage::ru_maxrss on Android, Linux and IOT, task_vm_info::phys_footprint on iOS and MacBook, and process_memory_counters::PrivateUsage on Windows. With the exception of task_vm_info::phys_footprint, accelerator (GPU/NPU/TPU) memory is not included. Memory measurements were taken from the second load which reads from loading caches. Memory usage on the first load may vary.

Android

Device Backend Text latency             
Google Pixel 11 Pro TPU 8.3 ms
S26 Ultra CPU 27.1 ms
S26 Ultra GPU 25.9 ms
Device Backend Text CPU memory
Google Pixel 11 Pro TPU 112 MB
S26 Ultra CPU 334 MB
S26 Ultra GPU 333 MB

iOS

Device Backend Text latency             
iPhone 18 Pro CPU 41.8 ms
iPhone 18 Pro GPU 11.6 ms
Device Backend Text CPU/GPU memory
iPhone 18 Pro CPU 84 MB
iPhone 18 Pro GPU 85 MB

Linux

Device Backend Text latency             
Arm 2.3 & 2.8GHz CPU 105 ms
NVIDIA GeForce RTX 4090 GPU 7.6 ms
Device Backend Text CPU memory
Arm 2.3 & 2.8GHz CPU 310 MB
NVIDIA GeForce RTX 4090 GPU 528 MB

macOS

Device Backend Text latency             
MacBook Pro M5 CPU 31.3 ms
MacBook Pro M5 GPU 9.5 ms
Device Backend Text CPU/GPU memory
MacBook Pro M5 CPU 165 MB
MacBook Pro M5 GPU 233 MB

Windows

Device Backend Text latency             
Dell XPS 16 (Intel Core Ultra Series 3) CPU 71.7 ms
Dell XPS 16 (Intel Core Ultra Series 3) GPU 19.2 ms
Dell XPS 16 (Intel Core Ultra Series 3) Running Intel OpenVINO NPU 13.3 ms
Device Backend Text CPU memory
Dell XPS 16 (Intel Core Ultra Series 3) CPU 211 MB
Dell XPS 16 (Intel Core Ultra Series 3) GPU 822 MB
Dell XPS 16 (Intel Core Ultra Series 3) Running Intel OpenVINO NPU 201 MB

Web

Device Backend Text latency             
MacBook Pro M5 GPU 21.8 ms

IoT

Device Backend Text latency             
Raspberry Pi 5 16GB CPU 161 ms
Jetson Orin Nano CPU 207 ms
Jetson Orin Nano GPU 88 ms
Arduino VENTUNO Q NPU 13.6 ms
Device Backend Text CPU memory
Raspberry Pi 5 16GB CPU 282 MB
Jetson Orin Nano CPU 302 MB
Jetson Orin Nano GPU 523 MB
Arduino VENTUNO Q NPU 206 MB
Downloads last month
3,155
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for litert-community/embeddinggemma-2-text-270m-litert-lm

Finetuned
(35)
this model

Space using litert-community/embeddinggemma-2-text-270m-litert-lm 1

Collection including litert-community/embeddinggemma-2-text-270m-litert-lm