speecht5_tts-pld-fil

microsoft/speecht5_tts finetuned on sapinsapin/pld.

Trained for 1000 steps on 1913 clips (batch 4×8, lr 1e-05, fp32 + gradient checkpointing). Synthesized listen-test samples are in samples/ (speechbrain x-vector speaker conditioning + microsoft/speecht5_hifigan vocoder).

metric value
eval_loss 0.4433

Trained with finetune_tts.py from the halohalo pipeline; the dataset adapter normalizes each corpus to (audio@16k, text, speaker_id) so corpora are swappable with a --dataset flag.

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sapinsapin/speecht5_tts-pld-fil

Finetuned
(1379)
this model