Is vision input officially supported?

#4
by you-n-g - opened

Hi, thanks for releasing XYZ-Aquila-mini!
I noticed that the model config includes a vision encoder and image/video tokens, and the repository is tagged as image-text-to-text. However, the model card and examples currently focus on text-only agentic search.

Should this checkpoint be considered officially capable of handling image or video inputs, or is text-only usage currently recommended?

Thanks!

Sign up or log in to comment