the modal is tiny but kv cache exploding!

#10
by rosspanda0 - opened

!

Nanbeige LLM Lab org

Thanks for bringing this up.
We have also investigated KV-cache sharing across loop passes, but the performance gains were notably smaller than with the full looped setup.
In future versions, we plan to mitigate the KV-cache overhead through linear and sparse attention mechanisms.

awesome!

leran1995 changed discussion status to closed

Sign up or log in to comment