Where can I deploy this model for inference?

by catworld1212 - opened Apr 27

Apr 27

Hi, I'm impressed with the work on InternVL and I'm interested in deploying its inference as an endpoint. Unfortunately, vLLM and TGI don't support this. Could anyone offer guidance on how to achieve this? I'd appreciate any suggestions you may have.

whai362

Apr 28

See "Chat Web Demo" at https://github.com/OpenGVLab/InternVL/blob/main/README.md

catworld1212

Apr 30

See "Chat Web Demo" at https://github.com/OpenGVLab/InternVL/blob/main/README.md

I want to deploy it as an inference not run it as a demo, Can you tell do InternVL-Chat-V1-5 requires flash attention?

catworld1212

May 2

Hi @whai362 @czczup what's the proper way to few-shot prompting (also called in-context learning? How do I give the previous context? I'm using lmdeploy to serve the inference can you help me, please?

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment