Running on <4gb Vram

#16
by Jit2024 - opened

Hi, I found out a way of running this model on <4gb Vram. And I have not used any quantization.
Repo: https://github.com/Jit-Roy/WeeLLM

Bro, nice method. But you shouldn't do spam across all available quantized models in HF

YarvixPA changed discussion status to closed

Sign up or log in to comment