Image-Text-to-Text
Transformers
Safetensors
English
qwen2_5vl_ca
feature-extraction
conversational
custom_code

This repository contains the model weights for CASA-Qwen2_5-VL-3B-Shared, introduced in the paper CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion. This is a variant of CASA-Qwen2_5-VL-3B where the self-attention and cross-attention layers share the same parameters.

See CASA-Qwen2_5-VL-3B for more information

Downloads last month
23
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kyutai/CASA-Qwen2_5-VL-3B-Shared

Finetuned
(1)
this model

Datasets used to train kyutai/CASA-Qwen2_5-VL-3B-Shared

Collection including kyutai/CASA-Qwen2_5-VL-3B-Shared

Paper for kyutai/CASA-Qwen2_5-VL-3B-Shared