May I as

#1
by wh-wh-wh - opened

May I ask, how did you achieve that? Yes, based on my practice, this architecture seems to require a base (dense) + MoE extraction layer. Are you using the MoE version — the community-improved 27B-a-17B?

I layered the qwen model myself to then rate layer activation on coding and reasoning prompts!
But wait I am currently working on making qwen 3.8 27B a 17BA MoE, is this something that already exists ?

Sign up or log in to comment