VideoVAP โ€” AVCocktail

Video-only stereo transformer Voice Activity Projection (VAP) turn-taking model.

Training data: Trained from scratch on the AVCocktail corpus.

Checkpoint: Final checkpoint (last.ckpt).

Files

  • hparams.yaml
  • checkpoint.ckpt

Source code

This checkpoint was produced with mm-turn-taking.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Collection including lggvu/videovap-avcocktail