Instructions to use laion/chatterbox-s3gen-vc-grow-ce-continue with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Chatterbox
How to use laion/chatterbox-s3gen-vc-grow-ce-continue with Chatterbox:
# pip install chatterbox-tts import torchaudio as ta from chatterbox.tts import ChatterboxTTS model = ChatterboxTTS.from_pretrained(device="cuda") text = "Ezreal and Jinx teamed up with Ahri, Yasuo, and Teemo to take down the enemy's Nexus in an epic late-game pentakill." wav = model.generate(text) ta.save("test-1.wav", wav, model.sr) # If you want to synthesize with a different voice, specify the audio prompt AUDIO_PROMPT_PATH="YOUR_FILE.wav" wav = model.generate(text, audio_prompt_path=AUDIO_PROMPT_PATH) ta.save("test-2.wav", wav, model.sr) - Notebooks
- Google Colab
- Kaggle
Chatterbox S3Gen VC · Continued rank-128 LoRA, 50/50 CE and emotion
This is a research-stage audio-to-audio voice-conversion adapter, not a standalone TTS model. It changes Chatterbox's S3Gen acoustic decoder while keeping the S3 semantic tokenizer, CAMPPlus speaker encoder and HiFT vocoder frozen. The adapter is rank 128 on 224 q/k/v/out attention projections, with the previously trained 99-score conditioning projector. Download the pinned upstream Chatterbox S3Gen and the 300-hour supervised/v1 RL parent release separately. This repository does not contain the upstream base weights or the training audio.
What this arm does
The completed v1 continue rank-128 adapter at step 4,000 is continued without merging it into the base.
The reward uses the frozen
Humaneness Ears Medium
model at revision 818506970d93c809a3295a0a02c40dee8ff2bfc3.
Content Enjoyment is specifically its audiobox:CE output, not the
separate eiv_extra:score_content_enjoyment output. For each 16-candidate
group, all reward components are standardized separately within that group;
the weighted advantage is clipped to ±2.5 and centered. The two emotion
dimensions are the strongest Ears emotions in the actual source recording,
not fixed labels from the prompt. The weights are 50% Content Enjoyment and 50% negative top-two emotion MSE.
The selected 4,000 source/target conditions contain 3,357 distinct sources
and 477 target-voice clips. They are traversed six times with different
permutations and sampling seeds: 24,000 optimizer steps and 384,000 generated
candidates per arm, but not 24,000 distinct input pairs. Every step is one
16-candidate group, split four candidates per GH200 GPU across a four-GPU
node, with synchronized gradients. The new AdamW schedule has 1,200 warm-up
steps (5%), a 2e-6 peak LR, cosine decay to 2e-7, and a frozen-policy flow
anchor of 0.025 to the respective v1 step-4,000 starting model. Checkpoints
with optimizer state are written every 1,000 groups and automatically
published under checkpoints/. See
RL_CHECKPOINTS.json for verified SHA-256, size and
step. A missing step means it has not yet been verified and published.
Inference and continuation
This is a research implementation for prepared Parquet rows, not a
one-command arbitrary-WAV VC pipeline. The source row supplies S3 semantic
tokens and dynamic emotion/style scores. A 5–10-second target reference row
supplies its acoustic prefix, CAMPPlus identity vector and static voice traits.
See TRAINING.md for preparation, exact model pins, score
order, training/resume commands and known limitations. The included
code/grow_followup/infer_followup.py
uses a prepared source/target pair:
pip install -r requirements.txt # install a compatible CUDA PyTorch first
hf download ResembleAI/chatterbox s3gen.safetensors \
--revision 5bb1f6ee58e50c3b8d408bc82a6d3740c2db6e18 \
--local-dir upstream_chatterbox
hf download laion/chatterbox-s3gen-vc-grow \
checkpoints/base_300h_adapter.safetensors \
--local-dir parent_v1
export CHATTERBOX_S3GEN_WEIGHTS="$PWD/upstream_chatterbox/s3gen.safetensors"
export CHATTERBOX_BASE_ADAPTER="$PWD/parent_v1/checkpoints/base_300h_adapter.safetensors"
export PYTHONPATH="$PWD/code:${PYTHONPATH:-}"
python code/grow_followup/infer_followup.py \
--pool /path/to/prepared_pool.parquet \
--score-order-json score_order_and_preparation.json \
--source-uid SOURCE_UID --target-uid TARGET_UID \
--arm continue --checkpoint checkpoints/step-0001000.pt \
--output-wav example.wav --seed 1901
The .pt files include optimizer and RNG state and use Python pickle;
load only trusted files. The inference code uses PyTorch's restricted
weights_only=True loader but still requires verified provenance. The
checkpoint is not a merged base model. Load the supervised base adapter, then load this continuation adapter.
Interpretation and limits
The rewards are model-predicted scores, not human ratings. Repeated traversal of 4,000 pairs can overfit, and the reward models can be exploited. The speaker cosine, when present, is a CAMPPlus proxy and cannot establish human-perceived identity fidelity. There is no claim here that reward improvement means better voice conversion. A held-out, matched-seed listening study with content accuracy, speaker similarity and independent naturalness metrics is required before selecting a default adapter. The selected sources are UID-disjoint from the previous holdout, not guaranteed speaker-disjoint.
LAION-authored adapter weights, code and documentation here are CC BY 4.0. The upstream Chatterbox dependency retains its separate MIT license and attribution. Use only voices and audio you have rights to process.
Model tree for laion/chatterbox-s3gen-vc-grow-ce-continue
Base model
ResembleAI/chatterbox