Agent Voice โ€” Piper ONNX

Agent Voice is a single-speaker, US English Piper/VITS voice for interactive agents. This release includes the standard combined ONNX model and the encoder/decoder graphs used by Piper's streaming synthesis path. It is the 24,000-update checkpoint selected by listening for the Cortexist voice classroom; it is not an LLM or an automatic speech recognition model.

Listen to a generated sample.

File Purpose
en_US-agent-voice.onnx Complete voice model for standard Piper synthesis
en_US-agent-voice.onnx.json Piper configuration, required beside the complete model
en_US-agent-voice.enc.onnx Text encoder, duration predictor, and flow for streaming
en_US-agent-voice.dec.onnx Chunked waveform decoder for streaming
SHA256SUMS Checksums for the uploaded files

Keep all four model files in the same directory. The paired .enc.onnx and .dec.onnx files are derived from the complete model; they are not separately trained voices. Audio is mono, 22,050 Hz, and uses eSpeak NG en-us phonemization. The model has one speaker.

Use

Standard Piper can synthesize from the combined model and its adjacent JSON:

piper --model en_US-agent-voice.onnx --output-file hello.wav "Hello from Agent Voice."

Streaming synthesis requires the implementation in Piper pull request #302, which is open as of this release. Until that code is available in your Piper installation, the split files alone will not enable --stream. With a Piper build containing the PR:

piper --model en_US-agent-voice.onnx --output-raw --stream \
  "Hello from Agent Voice." > hello.s16le

The raw output is signed 16-bit mono PCM at 22,050 Hz. For Python applications, load the same combined path with PiperVoice.load("en_US-agent-voice.onnx", streaming=True) and iterate over synthesize_stream(text); Piper finds the adjacent split graphs automatically.

The split encoder processes a whole sentence first, retaining its context and phoneme durations. The decoder then emits overlapping waveform chunks, cropping their edges before playback. This can reduce time to first audio, but it does not accept partial text or change the trained voice weights. Streaming omits whole-sentence peak normalization; use application or system gain as needed. The PR explains the implementation and its separate latency measurements.

Training and validation

The voice was fine-tuned from the Piper Joe medium voice. The training corpus began with 1,177 recordings: speech synthesized with VibeVoice-Large from a cleared voice reference, plus six retained original excerpts. After the project's transcript and audio review, 988 recordings (about 1.70 hours) were retained for this training run. The original reference and training recordings are not included in this repository. The project owner confirmed clearance to release the trained voice.

The selected checkpoint completed 24,000 paired generator/discriminator updates. The complete ONNX file is an exact copy of the validated export; the split graphs were generated from it. Local checks covered PyTorch/ONNX numerical agreement, combined/split agreement, chunked synthesis, phoneme timing, empty input and the Piper streaming output protocol. These checks establish export compatibility on the tested runtime, not a universal voice-quality rating or a speed guarantee.

The Joe source voice's model card identifies its dataset as CC0. That source-dataset label is not a license claim for this fine-tuned release. No downstream license for these fine-tuned weights is specified in this card.

This voice is a research release. Its naming does not imply affiliation with any fictional character, performer, or rights holder.

Downloads last month
33
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for cortexist/agent-voice

Finetuned
(4)
this model