speecht5_tts-fsc

microsoft/speecht5_tts finetuned for Filipino (Tagalog/Taglish) on sapinsapin/filipinospeechcorpus.

Trained for 1000 steps on 1867 read-speech clips (batch 4x8, lr 1e-05, fp32 + gradient checkpointing). Synthesized listen-test samples are in samples/ (speechbrain x-vector speaker conditioning + microsoft/speecht5_hifigan vocoder).

metric value
eval_loss 0.4432

Trained with finetune_tts.py from the halohalo pipeline; the dataset adapter normalizes each corpus to (audio@16k, text, speaker_id) so corpora are swappable with a --dataset flag.

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sapinsapin/speecht5_tts-fsc

Finetuned
(1372)
this model

Dataset used to train sapinsapin/speecht5_tts-fsc