FP8

#2
by vcerny - opened

Hi, great stuff, congrats! Will this works for FP8 quant too or does it require a new speculator? Thanks.

Yes it works with the FP8 quant target model. We tested it and the acceptance length is very close to the BF16 target model.

I tried running DFlash2 with unsloth's NVFP4 27B and vllm, from the pr mentioned, rejected it due to the quantized LM Head. Are you aware of this issue? Are you working on this, or might something else be the issue that I should verify? Am booting FP8 as I write this, but the extra KV headroom from the model size delta is something I'd ideally like to keep, even if sacrificing a little acceptance (MTP on the nvfp4 was in the 3-4 range during normal use, so it's still high quality).
Hardware is a DGX Spark.

Sign up or log in to comment