Instructions to use VohoAI/voho-saudi-speak-0.6b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use VohoAI/voho-saudi-speak-0.6b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="VohoAI/voho-saudi-speak-0.6b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("VohoAI/voho-saudi-speak-0.6b") model = AutoModelForCausalLM.from_pretrained("VohoAI/voho-saudi-speak-0.6b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use VohoAI/voho-saudi-speak-0.6b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "VohoAI/voho-saudi-speak-0.6b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VohoAI/voho-saudi-speak-0.6b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/VohoAI/voho-saudi-speak-0.6b
- SGLang
How to use VohoAI/voho-saudi-speak-0.6b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "VohoAI/voho-saudi-speak-0.6b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VohoAI/voho-saudi-speak-0.6b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "VohoAI/voho-saudi-speak-0.6b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VohoAI/voho-saudi-speak-0.6b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use VohoAI/voho-saudi-speak-0.6b with Docker Model Runner:
docker model run hf.co/VohoAI/voho-saudi-speak-0.6b
Voho Saudi Speak 0.6B
Turns formal Arabic into Arabic the way Saudis actually say it, in Najdi, Hijazi or Khaleeji, from Voho.
A voice agent that reads out written Arabic sounds like a news bulletin. This small model sits between the text and the voice: give it a formal sentence and a dialect, and it returns what a person from Riyadh, Jeddah or the Eastern Province would say on a phone call. It is 0.6B parameters, small enough to run on a CPU or next to a speech model on one GPU.
Results
chrF++ against what a Saudi speaker actually said, on 1,000 held-out sentences (higher is better):
| Dialect | Sentences | Formal text unchanged | Qwen3-0.6B, untrained | Voho Saudi Speak 0.6B | Gemini 2.5 Flash |
|---|---|---|---|---|---|
| All test sentences | 1,000 | 61.8 | 52.2 | 75.7 | 68.7 |
| Najdi (Riyadh, central) | 462 | 62.2 | 51.5 | 77.0 | 70.6 |
| Hijazi (Jeddah, Makkah) | 257 | 65.8 | 57.7 | 75.8 | 69.1 |
| Khaleeji (Eastern Province) | 281 | 56.6 | 47.3 | 73.2 | 64.7 |
"Formal text unchanged" is the floor: dialect and formal Arabic share most of their letters, so leaving the sentence as it is already scores well. Gemini 2.5 Flash is a much larger model doing the same job with the same instruction, shown for scale. Read the comparison with care: this model was trained on SADA and learned its spelling and phrasing, which Gemini never saw, and the formal inputs were themselves written by Gemini. On this test set that favours the small model; it does not mean a 0.6B model is better at Saudi Arabic than Gemini in general.
Examples
| Dialect | Formal input | Voho Saudi Speak 0.6B |
|---|---|---|
| Khaleeji | الآن فقط فهمت قصدك وسر اهتمامك. | الحين بس فهمت قصدك وسر اهتمامك. |
| Najdi | والله لست أرى أي شيء، أقسم بالله، نعم نعم ارفعها أكثر قليلا. | والله ماني شايف أي شيء أقسم بالله إيه إيه أرفعها شوي أكثر |
| Hijazi | ما هي همومك يا شوماخر؟ | إيش همومك يا شوماخر؟ |
Sentences it never saw in training, from a bank or a clinic line:
| Dialect | Formal input | Voho Saudi Speak 0.6B |
|---|---|---|
| Najdi | أين أنت الآن؟ أريد أن أحجز موعداً غداً. | وينك الحين أبغى أحجز موعد بكرة |
| Najdi | لا أستطيع رفع الحد إلى أكثر من ألف ريال، يجب أن تزور الفرع. | ما أقدر أرفع الحد إلى أكثر من ألف ريال لازم تزور الفرع. |
| Hijazi | سنرسل لك رمز التحقق الآن، من فضلك أخبرني به. | بنرسلك رمز التحقق دحين من فضلك خبرني فيه |
| Najdi | رقم طلبك هو 48213 وسيصل خلال ثلاثة أيام. | رقم طلبك هو 48213 وبيوصل خلال ثلاث أيام. |
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "VohoAI/voho-saudi-speak-0.6b"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)
DIALECT = {"najdi": "النجدية", "hijazi": "الحجازية", "khaleeji": "الخليجية الشرقية"}
def saudi(text, dialect="najdi"):
prompt = f"أعد صياغة هذه الجملة باللهجة السعودية {DIALECT[dialect]} كما يقولها شخص في مكالمة، بدون أي شرح:\n{text}"
ids = tok.apply_chat_template([{"role": "user", "content": prompt}], add_generation_prompt=True,
enable_thinking=False, return_tensors="pt")
out = model.generate(ids, max_new_tokens=96, do_sample=False)
return tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True).strip()
print(saudi("أين أنت الآن؟ أريد أن أحجز موعداً غداً."))
Training
- Base model:
Qwen/Qwen3-0.6B(Apache 2.0), full fine-tune, 2 epochs, one NVIDIA L4 - Targets: real Saudi speech. 84,394 transcribed sentences from SADA (Najdi, Hijazi and Khaleeji speakers), cleaned of noise transcriptions, repeats and duplicates
- Inputs: a formal Modern Standard Arabic version of each sentence, written by Gemini 2.5 Flash from the Saudi original. Pairs where Gemini changed nothing were dropped
- Test set: SADA's own test split, never seen in training
Licence and intended use
Non-commercial. SADA is licensed CC BY-NC-SA 4.0, so this model is released under the same licence: research and non-commercial use, with attribution, under the same terms.
For production Saudi Arabic voice, use the Voho API or the LiveKit plugin.
Limitations
- Trained on television speech: it knows how Saudis talk in dramas and interviews better than how they talk to a bank.
- Its inputs were written by another model, so it learns to undo that model's style of formal Arabic best.
- Small model: it can drop or change details in long or complicated sentences, and occasionally swaps who does what ("هل تريد أن أحولك" came out as "تبي تحولك"). Check anything customer-facing.
- Writes without diacritics.
Citation
Please cite SADA when using this model:
Saudi Audio Dataset for Arabic (SADA), Saudi Data and AI Authority (SDAIA), 2022.
- Downloads last month
- 207