PL-Flow pretrained bundle
Model weights for the PL-Flow two-stage speech synthesis implementation. Licenses differ by component; see THIRD_PARTY.md.
Components
| Component | Function | Origin | Weight license |
|---|---|---|---|
s1 |
Phoneme-aligned prosody flow, QK normalization | Trained from random initialization, EMA | GPL-3.0-only |
s2 |
Duration-conditioned acoustic flow, QK normalization | Trained from random initialization, EMA | GPL-3.0-only |
aligner |
HuBERT/text forced alignment | Trained from random initialization | GPL-3.0-only |
vocoder |
32 kHz mono, 32-channel 50 Hz SQ-GAN codec | Trained from random initialization, EMA | GPL-3.0-only |
plbert |
Frozen phoneme ALBERT | Extracted from Kokoro | Apache-2.0 |
hubert |
Frozen semantic features | TencentGameMate Chinese HuBERT | MIT |
speaker |
Frozen WavLM/ECAPA speaker features | External checkpoint linked by Seed-TTS-Eval | Download and upstream terms |
Six components contain config.json and model.safetensors. The speaker directory
contains configuration and download instructions only. bundle.json pins SHA256
hashes, the external speaker download, the prosody-code interface and inference
settings.
Usage
Install the PL-Flow source package in your own environment, then:
hf download rcell233/pl-flow --local-dir artifacts/pretrained
Download wavlm_large_finetune.pth
from the model link in Seed-TTS-Eval
and save it as artifacts/pretrained/speaker/wavlm_large_finetune.pth.
PL-Flow verifies the checkpoint's SHA256 and extracts speaker features
automatically during synthesis.
pl-flow verify --models artifacts/pretrained
pl-flow synthesize --models artifacts/pretrained \
--reference-audio data/reference.wav --reference-text '这是参考音频的文本。' \
--text '你好,欢迎使用语音合成。' --language zh --output outputs/example.wav
Default sampling uses S1 8 Euler steps / CFG 2.1, S2 10 Euler steps / CFG
2.8, temperature 1.0 and CUDA BF16 autocast for the flow models. Guidance is
(1 + scale) * conditional - scale * unconditional. Audio encoders and the codec
run in FP32. Context limits: 512 S1 tokens, 1024 total S2 frames, 250 prompt frames.
S1 and S2 use EMA weights. The bundle supports inference and initialization for training. Resuming training requires a full training checkpoint. The ASR aligner and SQ codec can also be trained using the PL-Flow package.