YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
ποΈ Voice Cloning Web Application
A production-quality, CPU-optimized Voice Cloning Web Application built for Windows using Resemble AI's Chatterbox TTS and Gradio. The application works completely offline after the initial model download and ensures minimal RAM/CPU footprint.
π Project Structure
app.py: Main Gradio Web UI script.install.py: Installer script that automatically sets up the environment, installs dependencies, and downloads the model.requirements.txt: List of dependencies.download_model.py: Script responsible for downloading the Chatterbox TTS model.config.py: Project configuration and paths.utils.py: Audio processing (format conversion, silence removal, normalization).model_loader.py: Handles loading the TTS model using a memory-optimized Singleton pattern.generate.py: Executes the model inference with memory garbage collection.outputs/: Saved generated audios.uploads/: Processed uploads.cache/: Cached Chatterbox model.venv/: Virtual environment.
βοΈ Installation
- Open a Command Prompt or PowerShell terminal.
- Navigate to this directory.
- Run the installer:
This will create a virtual environment, install CPU-only PyTorch (to save space), install all dependencies, and pre-download the model.python install.py
π Running the App
To make launching as easy as possible, a launcher script is included.
Simply double-click the start.bat file in your folder, or run it from your terminal:
.\start.bat
This script will automatically:
- Install any missing frontend dependencies.
- Launch the FastAPI AI Backend in a background window.
- Start the React Frontend.
- Open the stunning new Voice Cloning Studio in your browser.
The browser will automatically open your stunning new Voice Cloning Studio at http://localhost:5173.
π οΈ Usage
- Upload Voice: Upload a WAV/MP3/FLAC file. It must be between 5 and 30 seconds. The app will automatically convert, normalize, and remove silence.
- Type Text: Enter the text you want the AI to speak.
- Generate: Click the generate button. Generation time will depend on your CPU capability.
- Download: The audio will appear on the right side and is automatically saved in the
outputs/folder.
β FAQ & Troubleshooting
- Model takes too long to load? The first run might take a while to initialize. Subsequent generations use the cached model in memory.
- Out of Memory Error? Ensure you don't have other heavy applications open. The script automatically uses
gc.collect()to free memory. - Audio Error? Ensure your reference audio is clear, contains only one speaker, and is not overly noisy.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support