Excited to test :)

#1
by Darkknight535 - opened

πŸ˜„

πŸ˜„

Q4_K_M and Q4_K_S shipped, vision verified to work with the BF16 mmproj. I am working on a fork for llama.cpp to support this model's deepseek sparse attention; attention is currently running at a GQA style quadratic memory complexity. The NIAH test verified retrieval up to 32k kv cache

@Darkknight535 what were your thoughts on it?

(Also @patrickbdevaney thanks for your work!)

πŸ˜„

Q4_K_M and Q4_K_S shipped, vision verified to work with the BF16 mmproj. I am working on a fork for llama.cpp to support this model's deepseek sparse attention; attention is currently running at a GQA style quadratic memory complexity. The NIAH test verified retrieval up to 32k kv cache

🫑

@Darkknight535 what were your thoughts on it?

(Also @patrickbdevaney thanks for your work!)

gonna test.

gonna test.

Sounds good, let me know how it goes!

Sign up or log in to comment