Pruned version?
#3
by stfred - opened
From the size I am guesing this is based on the unpruned version, any plans to also quantize the pruned one for even smaller GPUs / longer contexts?
I don’t think they've released any pruned or turbo version; this is currently the only version for both FL2VA and REF2VA.
I mean like the default comfy version is pruned (see filename, making it just 21GB at int8) so quantizing that further one would expect much smaller file size Q4 / Q5 quants to be possible
Trying to quantize a model that’s already been pruned can totally tank its quality, making it pretty bad at generating video. So if the official version drops later, we’ll still get Q4 and Q5 at the same file size.