T5Gemma 2 1152-wide text encoder for vLLM

This checkpoint contains only the text encoder from google/t5gemma-2-1b-1b, mapped without weight changes to the native Gemma3ForCausalLM layout used by vLLM. Vision, decoder, and multimodal-only EOI embedding weights are omitted.

The configuration sets is_causal=false and use_bidirectional_attention=true, causing vLLM's Gemma3 implementation to use EncoderOnlyAttention. It has 26 layers, hidden size 1152, and 4/1 attention/KV heads. Weights and activations are BF16.

vLLM

vllm serve <model-id-or-path> --runner generate --no-enable-prefix-caching --max-model-len 32768

--no-enable-prefix-caching is required on vLLM 0.24 because encoder-only attention has no autoregressive KV cache.

Intended use

This is an infrastructure conversion for text encoding, retrieval, reranking, and experimentation. The tied LM head allows the vLLM generation runner to load the model, but this isolated encoder was not trained as a standalone causal language model.

The original Gemma terms apply. See the source model card.

Downloads last month
-
Safetensors
Model size
1.0B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jadanovitch/t5gemma-2-1b-encoder

Finetuned
(12)
this model