VideoVAP โ€” Candor

Video-only stereo transformer Voice Activity Projection (VAP) turn-taking model.

Training data: Trained from scratch on the Candor corpus.

Checkpoint: Fold 0, epoch 10.

Files

  • params.yaml
  • weights.pt

Source code

This checkpoint was produced with mm-turn-taking.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Collection including lggvu/videovap-candor