GLM 5.3 Flash Lite idea

#13
by ChessVania - opened

24b, 12b active (because 3b is fast, but bad)

alternatively, 60B-a6b

12b active is too much for a 24b parameters model.
5-6b would be the ideal.

alternatively, 60B-a6b

too big for most GPU's

12b active is too much for a 24b parameters model.
5-6b would be the ideal.

But much smarter than 6b

alternatively, 60B-a6b

too big for most GPU's

Best use of MoE models is to offload the experts to RAM. The active 6B parameters can fit in a 8GB GPU as long as you have enough RAM for the rest.

If they wanted they would, but since they don't, they won't.

YES PLEASE GLM 5.3 Flash Lite 24B A6B I WOUld Love that Please Z.AI You can then replace my Qwen3.8

I want 14-16B dense or 20-24B A5B MoE

I want 14-16B dense or 20-24B A5B MoE

Same I have 5070 ti

WE NEED A <10B AGAIN LIKE GLM 4 9B

70B dense or like a 150B MoE

70B dense or like a 150B MoE

Please don't that doesent fit on my gpu

Sign up or log in to comment