Gooya Bozorg v1.5

Gooya Bozorg v1.5 is the versioned PersianASR distribution of Thomcles/Chatterbox-TTS-Persian-Farsi, a Persian fine-tune of Resemble AI's 500M multilingual Chatterbox TTS model.

This release preserves the selected upstream checkpoint exactly. We did not claim additional training or altered weights: it was promoted after a controlled listening comparison against our own 227,753-clip step-3500 fine-tune, where the upstream Persian checkpoint retained slightly better pronunciation and pacing. The weight hashes are recorded in MODEL_MANIFEST.json.

Why “Bozorg”?

This is the larger, higher-quality Gooya lane. It prioritizes Persian speech quality and zero-shot voice cloning over the footprint of Gooya's small/on-device models.

Quick start

Install the official Chatterbox repository and the small runtime dependencies:

git clone https://github.com/resemble-ai/chatterbox.git
cd chatterbox
pip install -e .
pip install huggingface_hub safetensors soundfile

Then download this repository and run:

python inference.py \
  --model-dir /path/to/Gooya-Bozorg-v1.5 \
  --text "سلام، حالت چطوره؟" \
  --reference /path/to/reference.wav \
  --output gooya-bozorg.wav

The reference should be a clean single-speaker WAV. The defaults match the comparison run: exaggeration=0.5 and cfg_weight=0.5.

Included comparison samples

  • samples/canonical.wav: conversational Persian pronunciation
  • samples/conversational.wav: punctuation and dialogue-like pacing
  • samples/codeswitch.wav: mixed Persian and English terms

All three were generated with fixed seeds and the same reference/settings used for the competing Gooya checkpoint.

Provenance

  • Upstream model: Thomcles/Chatterbox-TTS-Persian-Farsi
  • Pinned upstream snapshot: 4e9f6b7043d30d8328bb842b6126676bead8d9de
  • Architecture/runtime: ResembleAI/chatterbox
  • Weight status: byte-identical repackaging of the selected upstream files
  • PersianASR release name: Gooya Bozorg v1.5

Credit for the Persian training and original release belongs to Thomcles. Credit for Chatterbox belongs to Resemble AI.

Limitations

  • This model can still mispronounce ambiguous Persian words and mixed-script text.
  • Pacing and question intonation are improved relative to our tested fine-tune, but are not fully controllable.
  • Output quality depends strongly on the reference clip.
  • This release is not a tiny/on-device model and needs substantially more memory than Gooya's small models.
  • The license is non-commercial; review it before deployment.

License

The upstream checkpoint is released under CC BY-NC 4.0. This repository keeps the same license and attribution requirements. It is not licensed for commercial use.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Reza2kn/Gooya-Bozorg-v1.5

Finetuned
(1)
this model