Not sure if you are still working on this but
#3
by lainlives - opened
I have made a few voice models using 110000 during training, it actually very measurably increases the quality of inference in rmvpe/+ as well.
Not sure if you found what you were looking for with this model but I will claim this finetune's far more useful in training than inference.
Basically, all of the RMVPE, DJCM, and HPA variants in my repository are attempts to improve upon the original models. I'm never completely sure whether those improvements consistently hold up in practice, so I always consider them experimental. I'll probably keep experimenting if I come up with a new architecture or approach.