MM-Turn-Taking-Candor-AVCocktail
Collection
Pretrained models related to Candor & AVCocktail datasets. โข 11 items โข Updated
Audio-visual early-fusion Voice Activity Projection (VAP) turn-taking model.
Training data: Trained from scratch on the AVCocktail corpus.
Checkpoint: Epoch 5, val/loss=3.3068 (lowest val loss checkpoint; note the run's epoch 3 checkpoint was not retained on disk).
hparams.yamlcheckpoint.ckptThis checkpoint was produced with mm-turn-taking.