Instructions to use jadanovitch/t5gemma-2-1b-encoder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jadanovitch/t5gemma-2-1b-encoder with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="jadanovitch/t5gemma-2-1b-encoder")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("jadanovitch/t5gemma-2-1b-encoder") model = AutoModelForCausalLM.from_pretrained("jadanovitch/t5gemma-2-1b-encoder", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jadanovitch/t5gemma-2-1b-encoder with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jadanovitch/t5gemma-2-1b-encoder" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jadanovitch/t5gemma-2-1b-encoder", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/jadanovitch/t5gemma-2-1b-encoder
- SGLang
How to use jadanovitch/t5gemma-2-1b-encoder with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jadanovitch/t5gemma-2-1b-encoder" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jadanovitch/t5gemma-2-1b-encoder", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jadanovitch/t5gemma-2-1b-encoder" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jadanovitch/t5gemma-2-1b-encoder", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use jadanovitch/t5gemma-2-1b-encoder with Docker Model Runner:
docker model run hf.co/jadanovitch/t5gemma-2-1b-encoder
T5Gemma 2 1152-wide text encoder for vLLM
This checkpoint contains only the text encoder from google/t5gemma-2-1b-1b,
mapped without weight changes to the native Gemma3ForCausalLM layout used by
vLLM. Vision, decoder, and multimodal-only EOI embedding weights are omitted.
The configuration sets is_causal=false and
use_bidirectional_attention=true, causing vLLM's Gemma3 implementation to use
EncoderOnlyAttention. It has 26 layers, hidden size
1152, and 4/1
attention/KV heads. Weights and activations are BF16.
vLLM
vllm serve <model-id-or-path> --runner generate --no-enable-prefix-caching --max-model-len 32768
--no-enable-prefix-caching is required on vLLM 0.24 because encoder-only
attention has no autoregressive KV cache.
Intended use
This is an infrastructure conversion for text encoding, retrieval, reranking, and experimentation. The tied LM head allows the vLLM generation runner to load the model, but this isolated encoder was not trained as a standalone causal language model.
The original Gemma terms apply. See the source model card.
- Downloads last month
- -
Model tree for jadanovitch/t5gemma-2-1b-encoder
Base model
google/t5gemma-2-1b-1b