audioforge heads

Small trained heads for audioforge, a real-time speech front end for voice agents. They sit on top of one frozen NVIDIA streaming model and add voice activity, "is it the user?" and "is the turn over?". These files are the heads only. The NVIDIA models are not redistributed here; audioforge-download fetches them from NVIDIA under their own licences.

file core what it is
served_heads_v0.4.pt 115M (default) VAD, turn, speaker and speech heads for stt_en_fastconformer_hybrid_large_streaming_multi
served_heads_0p6b_v0.4.pt 0.6B the same heads retrained on nemotron-speech-streaming-en-0.6b
tsvad_spk.pt, tsvad_0p6b.pt 115M, 0.6B target-speaker VAD head
voice_gender_115m.pt, voice_gender_0p6b.pt 115M, 0.6B perceived voice gender head

Use

git clone https://github.com/maxmelichov/audioforge && cd audioforge
uv run audioforge-download   # fetches the NVIDIA model and checks these heads
uv run audioforge-serve

The same files ship in the GitHub repo's assets/, so the download works without this page. Results and the test protocol are in the repo's docs/RESULTS.md.

Licence

The heads are Apache-2.0. The NVIDIA base models keep their own licences: CC-BY-4.0 for the 115M model (credit to NVIDIA), NVIDIA Open Model License for the 0.6B model.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for notmax123/audioforge-heads

Finetuned
(15)
this model