This is a Gemma3 model uploaded using the KerasHub library and can be used with JAX, TensorFlow, and PyTorch backends. This model is related to a CausalLM task.

Model config:

  • name: gemma3_backbone
  • trainable: True
  • dtype: {'module': 'keras', 'class_name': 'DTypePolicy', 'config': {'name': 'float32'}, 'registered_name': None}
  • vocabulary_size: 262144
  • image_size: 896
  • num_layers: 34
  • num_query_heads: 8
  • num_key_value_heads: 4
  • hidden_dim: 2560
  • intermediate_dim: 10240
  • head_dim: 256
  • query_head_dim_normalize: True
  • use_query_key_norm: True
  • use_post_ffw_norm: True
  • use_post_attention_norm: True
  • attention_logit_soft_cap: None
  • final_logit_soft_cap: None
  • use_sliding_window_attention: True
  • sliding_window_size: 1024
  • local_rope_scaling_factor: 1.0
  • global_rope_scaling_factor: 8.0
  • vision_encoder: {'module': 'keras_hub.src.models.gemma3.gemma3_vision_encoder', 'class_name': 'Gemma3VisionEncoder', 'config': {'name': 'gemma3_vision_encoder', 'trainable': False, 'dtype': {'module': 'keras', 'class_name': 'DTypePolicy', 'config': {'name': 'float32'}, 'registered_name': None}, 'num_heads': 16, 'hidden_dim': 1152, 'num_layers': 27, 'intermediate_dim': 4304, 'output_dim': 2560, 'pool_size': 4, 'image_size': 896, 'patch_size': 14, 'layer_norm_epsilon': 1e-06}, 'registered_name': 'keras_hub>Gemma3VisionEncoder'}
  • use_bidirectional_attention: False
  • layer_norm_epsilon: 1e-06
  • dropout: 0
  • is_embedding_model: False
  • pooling_intermediate_dim: None
  • embedding_dim: None

This model card has been generated automatically and should be completed by the model author. See Model Cards documentation for more information.

Downloads last month
79
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support