Gooya Koochik v2.0-exp — Q6-TractCUDA

Standalone Persian TTS bundle for the native Rust Shenava tract CUDA fork. The model uses symmetric grouped W6A32 speech weights (6-bit, group size 32), stored in an int8 container with Zstandard compression and dequantized to FP32 arithmetic. Frontend, codec, embeddings, and output heads remain FP32.

The compressed bundle is about 504 MB. This is a native Rust runtime artifact; Python is used only to launch the compiled binary.

One-line run

import subprocess; subprocess.run(["./gooya_koochik", "./Q6-TractCUDA", "output.wav", "سلام، حالت چطوره؟"], check=True)

Build the launcher from gooya-bozorg-native, with the Shenava tract CUDA dependency enabled, and run on a CUDA system with the required CUDA libraries. Set GOOYA_KOOCHIK_DEVICE=cuda to select the backend.

Validation

The bundle passed the finite 12-phrase Shenava transcription parity suite at 98.507% (1 differing word out of 67). This is a transcription gate, not a claim of perceptual quality or unseen-text accuracy. The CUDA backend was validated on an RTX 5080 using the private Shenava tract fork.

Source checkpoint: Reza2kn/Gooya-Koochik-v2.0-exp revision 537bb48320fd657415eef5b4fb6796b6cebab195.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Reza2kn/GooyaKoochikv2-exp-Q6-TractCUDA

Finetuned
Qwen/Qwen3-0.6B
Finetuned
k2-fsa/OmniVoice
Quantized
(2)
this model