MMVAP โ€” AVCocktail

Audio-visual early-fusion Voice Activity Projection (VAP) turn-taking model.

Training data: Trained from scratch on the AVCocktail corpus.

Checkpoint: Epoch 5, val/loss=3.3068 (lowest val loss checkpoint; note the run's epoch 3 checkpoint was not retained on disk).

Files

  • hparams.yaml
  • checkpoint.ckpt

Source code

This checkpoint was produced with mm-turn-taking.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Collection including lggvu/mmvap-avcocktail