Beats native MTP head by 10-20% across the board

#1
by zxbc2023 - opened

The quantization targeting here is very effective, I am running mostly Q4 and Q5 models and this dspark head is winning against all built-in MTP heads.

Also, I'd like to point out that your model beats magnitudedev's Dspark head, by quite a margin.

A trade off for VRAM, for sure. I am interested in Q6 or Q4 quantizations' relative performance now. I will report back to let you know whether they represent useful alternatives.

Yeah man the draft model seems to generalize well. I'm interested in Q6 to save a bit more space for mmprojector or context. If it works we can potentially remove the vendor mtp head completely for further saving.

Thanks everyone for trying it out! We'll keep training new versions to improve the acceptance rate in agent scenarios β€” this version was actually trained on just 1/10 of our data.

Sign up or log in to comment