We don't need BF16 GGUF, we need Q8 or Q4 quant weight

#1
by weicj - opened

it doesn't make sense for us to use the BF16 gguf at this size, we need to fit this model into some 6GB GPU

You can check my repo.

Sign up or log in to comment