EmbeddingGemma-2 0.8 BPW Multimodal (LittleBit Quantized)

EmbeddingGemma-2 0.8 BPW Multimodal is an ultra-compact, sub-1-bit quantized runtime derivative of google/embeddinggemma-2 developed for extreme memory efficiency across Text, Image, Video, and Audio in a single unified 768-dimensional embedding space.

Following Google's full 740M parameter architecture (270M Text Backbone + 170M Vision/Video Encoder + 300M Audio Encoder), this model supports cross-modal search across any combination of text, images, video clips, and audio recordings. By applying Samsung Research's LittleBit sub-1-bit quantization framework ([arXiv:2506.13771]) with an asymmetric 65/35 latent factorization across all three encoders, weights are compressed to 0.8 Bits-Per-Weight (BPW).

This slashes runtime RAM from ~2,960 MB (BF16/FP32 full base) down to ~704.0 MB (a 76.2% reduction in model memory, or $\sim 4.2\times$ smaller).

Base Model & Derivative Notice

  • Base Model: google/embeddinggemma-2
  • Modality: Omni-Modal (Text + Image + Video + Audio)
  • Original Model Developers: Google DeepMind
  • Quantized by: lethalbeats (LethalBeats)
  • Distribution Notice: This repository is a community-contributed quantized runtime artifact and is not an official Google distribution. It is subject to the Gemma Terms of Use.

Memory Footprint & Runtime Execution

Model weights reside in RAM strictly in packed bitstream format (~704.0 MB). Unpacking occurs on-the-fly inside CPU SIMD registers (ymm0, ymm1 in AVX2) in chunks of 8/16 elements per clock cycle with $< 64\text{ KB}$ working cache overhead.

Model Configuration Active Modalities Model Weights RAM Reduction vs Base
EmbeddingGemma-2 (Full Base) Text, Image, Video, Audio ~2,960 MB Baseline (0%)
EmbeddingGemma-2 0.8 BPW Multimodal (This Model) Text, Image, Video, Audio ~704.0 MB -76.2%

Note: In-register streaming dequantization prevents in-memory RAM spikes, allowing unified omni-modal search inside edge hardware, appliances, and micro-instances without GPU requirements.


Quickstart & Usage

Adheres to Google's official SentenceTransformers dictionary input convention:

import numpy as np
from modeling_littlebit import EmbeddingGemma2MultimodalLittleBit

# Load model from local directory or Hugging Face Hub
model = EmbeddingGemma2MultimodalLittleBit.from_pretrained("lethalbeats/embeddinggemma-2-0.8bpw-multimodal")

# 1. Text Query
text_emb = model.encode("Nature scene with thunderstorm and night sky.")

# 2. Image Embedding
image_emb = model.encode({"image": "aurora_mountain.jpg"})

# 3. Video Embedding (sampled at 1 fps via vision encoder)
video_emb = model.encode({"video": "storm_timelapse.mp4"})

# 4. Audio Embedding (16 kHz mono)
audio_emb = model.encode({"audio": "thunderstorm.wav"})

# Cross-modal similarities
sim_image = float(text_emb[0] @ image_emb[0])
sim_video = float(text_emb[0] @ video_emb[0])
sim_audio = float(text_emb[0] @ audio_emb[0])

print(f"Similarity (Text vs Image): {sim_image:.4f}")
print(f"Similarity (Text vs Video): {sim_video:.4f}")
print(f"Similarity (Text vs Audio): {sim_audio:.4f}")

Matryoshka Dimension Slicing (MRL)

EmbeddingGemma-2 natively supports Matryoshka Representation Learning (MRL). You can slice embeddings to 512 or 256 dimensions and re-normalize for additional storage savings in vector databases:

# Truncate to 256 dimensions and re-normalize L2
emb_256 = text_emb[:, :256]
emb_256 = emb_256 / np.linalg.norm(emb_256, axis=-1, keepdims=True)
print("Sliced embedding shape:", emb_256.shape)

Model Specifications

Parameter Value
Base Architecture Gemma 2 Multimodal Omni-Transformer
Base Model google/embeddinggemma-2
Supported Modalities Text, Image, Video, Audio
Output Dimension 768 (supports MRL slicing to 512, 256)
Quantization Method LittleBit 0.8 BPW
Capacity Distribution Asymmetric 65% Primary / 35% Secondary
Model Weights RAM Footprint ~704.0 MB
Max Text Sequence Length 8,192 tokens
Video Processing Sampled at 1 fps via Vision Encoder
Audio Sample Rate 16,000 Hz mono
Similarity Function Cosine Similarity (Dot product on unit $\mathbb{S}^{767}$)

Theoretical Foundation: LittleBit (Samsung Research / ICML)

This model's quantization relies directly on the breakthroughs introduced by Samsung Research in the paper:

LittleBit: Ultra Low-Bit Quantization via Latent Factorization
Samsung Research (ICML)
arXiv:2506.13771 | HTML Full Paper (v5)


Citations & References

If you use this model in your research or applications, please cite:

@article{littlebit_2025,
  title={LittleBit: Ultra Low-Bit Quantization via Latent Factorization},
  author={Samsung Research},
  journal={arXiv preprint arXiv:2506.13771},
  year={2025},
  url={https://arxiv.org/abs/2506.13771}
}

@article{embedding_gemma_2025,
  title={EmbeddingGemma: Powerful and Lightweight Text Representations},
  author={Schechter Vera, Henrique and Dua, Sahil and Zhang, Biao and Salz, Daniel and Mullins, Ryan and Raghuram Panyam, Sindhu and Smoot, Sara and Naim, Iftekhar and Zou, Joe and Chen, Feiyang and Cer, Daniel and Lisak, Alice and Choi, Min and Gonzalez, Lucas and Sanseviero, Omar and Cameron, Glenn and Ballantyne, Ian and Black, Kat and Chen, Kaifeng and Wang, Weiyi and Li, Zhe and Martin, Scott},
  journal={arXiv preprint},
  year={2025}
}

@misc{embeddinggemma2_google,
  title={EmbeddingGemma-2: Multimodal Representation Models},
  author={Google DeepMind},
  year={2025},
  publisher={Hugging Face},
  howpublished={\url{https://huggingface.co/google/embeddinggemma-2}}
}

License

This model inherits the Gemma Terms of Use from Google.

Downloads last month
12
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lethalbeats/embeddinggemma-2-0.8bpw-multimodal

Quantized
(57)
this model

Paper for lethalbeats/embeddinggemma-2-0.8bpw-multimodal