ChiReSSD

Speaker-preserving reconstruction of disordered speech.

Given a target transcript and a short reference recording of a speaker, ChiReSSD synthesizes that transcript with canonical pronunciation while keeping the speaker's voice and prosody. Pronunciation enters through the text pathway; identity and prosody come from the style pathway. That separation is the point: ordinary style-preserving TTS treats disordered articulation as part of the speaker's style and so reproduces the mispronunciation it was meant to correct.

Model description

StyleTTS2 fine-tuned on child disordered speech. Two 128-dimensional style vectors are extracted from the reference recording โ€” acoustic (timbre) and prosodic โ€” and interpolated with a style sampled from the adapted diffusion prior. alpha weights the acoustic side and beta the prosodic; 1.0 is fully diffusion-sampled, 0.0 fully reference-driven.

It works on unseen speakers from a reference as short as a few seconds. No per-speaker model is trained, and no paired typical/atypical recordings are required.

Base model and license chain

Fine-tuned from the StyleTTS2 LibriTTS second-stage checkpoint by yl4579 (MIT). That base checkpoint is not redistributed here โ€” obtain it from the upstream release. The frozen helper models (ASR text aligner, JDCNet pitch extractor, PL-BERT) likewise ship inside the upstream repository and are not redistributed.

ChiReSSD builds upon StyleTTS2, with its modifications published as patch files in the code repository.

Intended use

Research on speech reconstruction and on automated clinical evaluation of speech sound disorders.

Out of scope

  • Not a medical device. No diagnostic or treatment decision should rest on its output.
  • Not a replacement for assessment by a licensed speech-language pathologist.

Training data

Child speech from UltraSuite/CLP-derived recordings, obtained under their own data use agreements. The corpus is not released here and is not redistributable: it is identifiable child clinical speech. To fine-tune your own model, https://huggingface.co/datasets/changelinglab/ultrasuite-benchmark is a suitable starting point; see DATA.md in the code repository.

Training configuration

Epochs 4
Batch size 4
Max length 600 frames
Learning rate 1e-5 (lr, bert_lr, ft_lr)
lambda_F0 20 (upstream: 1)
lambda_mel 5
Style diffusion from epoch 2
Joint SLM-adversarial from epoch 3
Decoder HiFi-GAN, multispeaker
Sample rate 24 kHz
LR schedule OneCycleLR, pct_start=0.1 (upstream: 0)

The heavy pitch weighting is deliberate: the pitch extractor was pretrained on adult voices, and children's F0 is both higher and more variable, so it needs the strongest adaptation of any component. Conversely, only four epochs โ€” longer schedules start fitting the disordered articulation itself.

Trained on 2ร— 48 GB GPUs. At batch 4 and max_len 600 the recipe needs more than 48 GB, so a single smaller card requires lowering both.

Inference

Two presets ship with the code:

Preset alpha beta steps Purpose
default 0.8 0.5 10 The released operating point
torgo 1.0 0.5 15 Adult dysarthric speech

alpha is high on purpose. A low alpha leans on the reference acoustics, which is exactly where the disordered articulation lives.

Synthesis is stochastic. The initial style latent is zeros rather than a Gaussian draw, but the ADPM2 sampler is ancestral and adds fresh noise at every step, so repeated calls differ in waveform and in duration. Pass a seed for reproducible output.

Usage

from chiressd.model import ChiReSSD
from chiressd.config import load_preset

model = ChiReSSD.from_pretrained()  # downloads this checkpoint
style = model.compute_style("speaker_reference.wav")
wav = model.synthesize(
    "butterfly butterfly butterfly",
    style,
    seed=1234,
    **load_preset("default").as_kwargs(),
)

Run chiressd-setup first: it clones and patches the upstream StyleTTS2 checkout that supplies the architecture and frozen helper models.

Checkpoint provenance

Derived from the fine-tuning run's final checkpoint by keeping state['net'] only, removing the module. prefix that DataParallel added to ten of the thirteen submodules, and making every tensor detached, CPU-resident and contiguous. Precision is unchanged (float32; no fp16 cast, which would alter outputs).

Citation

Rosero, Yeo, Mortensen, Van't Slot, Hallac, and Busso. Generative Reconstruction of Pediatric Disordered Speech for Automated Clinical Evaluation. In Proceedings of the IEEE Spoken Language Technology Workshop 2026 (SLT '26), Palermo, Italy, 2026.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support