Is this model's architecture already supported by llama.cpp? Also, what are the vram requirements to run this model locally?
1: I believe so, it is simply safetensors. 2: a couple terabytes nothing much.
· Sign up or log in to comment