New Higher quality Quant.

#2
by Harsha001 - opened

Hey, i just wanted to ask when will the Higher Quality Quant with the 15-20tps decode would be available with also a decent Vision encoder. If possible i'd like for you to also explore the n-gram ssd streaming with a MTP like predictor loaded for the ngram predictive loading onto the ram. Ik i am asking too much but the SSD ngram part is already supported i tested it, results in 12tps decode and that's bad, so was thinking an experienced developer can pull-off this n-gram predictor on the NPU that would make this model and similar "separate n-gram" architecture models a killer in performance and speed on STRIX devices.

SSD loading should be coming. I am also working on higher quality quants. My goal with the FP4 imatrix quant was to make something that easily fits in to memory on a strix halo and is still high quality. It beats some other quants as is explained in the readme.

The FP4 imtrax quant is very good. Vision tower is coming today hopefully. I created mainline compliant quants as well. Q4 and Q5. Enjoy!

Thanks, btw did you check with my proposal, tbh it's a Hypothesis to be tested.. ssd streaming but with a MTP like predictor for prefetching the next n-gram token.

The NGRAM lookup is really cheap and and a predictor won't do anything there. The cost is in the inference pipeline.

Sign up or log in to comment