Why is it much slower and much more memory hungry than slightly larger models such as Qwen3.5-4B?

#31
by slightlyoutofphase - opened

Definitely not what I was expecting.

Nanbeige LLM Lab org

Thank you for raising this. The Nanbeige4 series (Nanbeige4, Nanbeige4.1, Nanbeige4.2, and the upcoming Nanbeige4.5) focuses on improving model quality under a fixed parameter budget. In Nanbeige4.2, the Looped Transformer and disabled KV sharing improve model performance but result in slower decoding.

We are preparing a DFlash version of Nanbeige4.2 for faster inference. Nanbeige5 will use linear attention to further improve inference efficiency.

What's the ETA for Nanbeige 4.5?

Waiting for this DFlash . It's too slow for me right now. Just a half of gemma-4-e4b on my side.

Sign up or log in to comment