Compressing visual tokens in vision-language models: 3x more requests per GPU on Qwen2-VL Jun 17 • 1