audioforge heads
Small trained heads for audioforge, a real-time speech front end for
voice agents. They sit on top of one frozen NVIDIA streaming model and add voice activity, "is it the user?" and
"is the turn over?". These files are the heads only. The NVIDIA models are not redistributed here;
audioforge-download fetches them from NVIDIA under their own licences.
| file | core | what it is |
|---|---|---|
served_heads_v0.4.pt |
115M (default) | VAD, turn, speaker and speech heads for stt_en_fastconformer_hybrid_large_streaming_multi |
served_heads_0p6b_v0.4.pt |
0.6B | the same heads retrained on nemotron-speech-streaming-en-0.6b |
tsvad_spk.pt, tsvad_0p6b.pt |
115M, 0.6B | target-speaker VAD head |
voice_gender_115m.pt, voice_gender_0p6b.pt |
115M, 0.6B | perceived voice gender head |
Use
git clone https://github.com/maxmelichov/audioforge && cd audioforge
uv run audioforge-download # fetches the NVIDIA model and checks these heads
uv run audioforge-serve
The same files ship in the GitHub repo's assets/, so the download works without this page. Results and the test
protocol are in the repo's docs/RESULTS.md.
Licence
The heads are Apache-2.0. The NVIDIA base models keep their own licences: CC-BY-4.0 for the 115M model (credit to NVIDIA), NVIDIA Open Model License for the 0.6B model.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for notmax123/audioforge-heads
Base model
nvidia/nemotron-speech-streaming-en-0.6b