A possible middleway

#8
by mkristian - opened

This is pure speculation but I wonder whether you considered this approach and decided for the one you used. This model takes ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:IQ3_S and prunes 50% of the experts using some GCO way to decide which one to keep and which one to prune.

If you would start with ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF:Q2_0 and start pruning experts from there but let's say only 30-40% of the experts, just enough to have this part (!= n-gram one) fit into a 32GB GPU. Would such a 'middleway' perform better then this one ? Or is there some obvious reasons why this middleway will not be good ?

BTW big thanks for all the quants - very much appreciated.

IST Austria Distributed Algorithms and Systems Lab org

Thanks for the suggestion! The main issue is that the full Q2_0 model already gets only around 81% on LiveCodeBench v6, whereas this pruned version gets around 86%. So even before pruning, Q2_0 starts roughly 5 percentage points behind. Pruning another 30–40% of its experts would likely introduce some additional degradation, so I would not expect it to outperform the current approach.

In other words, for this model it seems preferable to keep fewer experts at a somewhat higher precision rather than retain more experts but quantize all of them down to Q2_0. There may still be a sweet spot somewhere between the two, and it would definitely be interesting to test a less aggressive pruning ratio, but based on the current LiveCodeBench results, starting from Q2_0 does not look particularly promising.

Sign up or log in to comment