it_IT-serena β Piper training checkpoint (Italian, medium)
Finetuning base checkpoint for Piper TTS (VITS), trained from scratch on Italian.
- Model: piper (piper_train), quality medium, 22.05 kHz
- Base voice: serena β female, neutral Standard Italian (no regional accent), synthetic
- Training: 94 epochs from scratch, batch 10, 29,470 clips (27 h)
- Dataset: committa/serena-synthetic-it-27h
- File:
epoch=94-step=559930.ckptβ keep this filename (piper_train convention:epoch=N-step=M), which tooling uses to pick the highest checkpoint
Why this checkpoint
Piper publishes no official pretrained checkpoint for Italian (the rhasspy/piper-checkpoints
catalog has no it entry; the only official Italian voices are it_IT-paola-medium, which was
finetuned from a U.S. English base, and it_IT-riccardo-x_low, trained from scratch at x_low
quality β neither exposes a trainable checkpoint). This is the only public Italian medium
checkpoint that can be used as a finetuning base.
Sample
Audio of the base voice (it_IT-serena-medium):
Use as a base for finetuning other Italian voices
Any checkpoint of the same quality config (medium) can be resumed with piper_train:
python -m piper_train \
--dataset-dir /path/to/your_voice_training_folder \
--resume_from_checkpoint epoch=94-step=559930.ckpt \
--quality medium
Because the base already speaks Italian (correct phonemes and prosody), a new Italian voice converges in far fewer epochs than from scratch β and with less data. This works for both female and male target voices: the timbre is learned from the new dataset, so with enough epochs the new voice fully takes over.
License
CC-BY-4.0 β derived from committa/serena-synthetic-it-27h. If you use this checkpoint to build a voice, credit the dataset and this model.