It's such a shame. I've tried Qwen3.8-27B, from base, to FP8, to several different NVFP4 models, as well as different permutations of chat templates on vLLM. They all go crazy and fail my benchmark. Probably the best NVFP4 model though.
DavidAU's NEO CODER model is the best tune so far.
· Sign up or log in to comment