Jev-Multilingual: Urdu, Sindhi & Roman Decisions for Python
A multilingual extension of muhammadnoman76/jev-urdu, adding native Sindhi (سنڌي) and Roman Sindhi decision-making capabilities to the Jev architecture.
Base model developed by Muhammad Noman through LughaatNLP.
Sindhi and Roman Sindhi adaptation, presets, and benchmarks developed by Shakeel Ahmed Sanjrani.
Model Fork | Upstream Base Model | Apache-2.0 license | Python 3.10+
From a message to a structured decision
Urdu and Sindhi in customer applications arrive in multiple scripts: Arabic script, Roman script (Latin phonetics), and English mixed text.
jev-multilingual provides a unified, typed decision model across both languages:
import jev_urdu
model = jev_urdu.load("shakeel143/jev-multilingual")
# 1. Native Sindhi Script Triage & Sentiment
sindhi_msg = "منهنجو يونيورسٽي پورٽل ٽن ڏينهن کان لاگ ان نٿو ٿئي، فوري مدد گهرجي."
print(model.triage(sindhi_msg, lang="sd"))
print(model.sentiment("هي سروس تمام بهترين ۽ سٺي آهي.", lang="sd"))
# 2. Native Sindhi Topic Classification
topic_msg = "اسان جي زمين ۾ ڪڻڪ جو فصل ڏاڍو سٺو ٿيو آهي."
print(model.topic(topic_msg, lang="sd"))
# 3. Roman Sindhi Decisions
print(model.sentiment("hee phone daadho sutho aa, battery zabardast aa.", lang="sd"))
# 4. Urdu (100% Backwards Compatible)
print(model.sentiment("یہ فون بہت اچھا ہے، بیٹری بھی زبردست ہے۔", lang="ur"))
The model returns structured decisions and calibrated probability distributions rather than conversational text generations.
What is new in Jev-Multilingual?
- Native Sindhi Task Presets (
SINDHI_PRESETS):- Pre-configured, natural Sindhi question prompts and label spaces for triage, sentiment, claim verification, and topic routing.
- Language-Aware API:
- Built-in
sentiment(),triage(),check_claim(), andtopic()methods supportlang="sd"for Sindhi andlang="ur"for Urdu.
- Built-in
- Roman Sindhi Evaluation & Lexical Grounding:
- Evaluated and adapted on authentic Roman Sindhi vocabulary (
daadho,sutho,ghandi,mayosi,bekaar).
- Evaluated and adapted on authentic Roman Sindhi vocabulary (
- Antonym Contradiction Resolution:
- Fine-tuned decision head resolves zero-shot antonym confusion in Sindhi (e.g.
گرم ↔ ٿڌي).
- Fine-tuned decision head resolves zero-shot antonym confusion in Sindhi (e.g.
- Zero Catastrophic Forgetting:
- Base Urdu benchmarks remain 100% shielded and regression-tested.
Built-in Tasks
| Method | Languages | Decisions Returned |
|---|---|---|
triage(text, lang="ur"|"sd") |
Urdu, Sindhi | Issue (بنيادي مسئلو), Human request (انساني مدد), Priority (ترجيح) |
sentiment(text, lang="ur"|"sd") |
Urdu, Sindhi, Roman | Sentiment (مثبت/منفي/غير جانبدار), Dissatisfaction |
topic(text, lang="sd") |
Sindhi, Roman Sindhi | Domain category (تعليم, ٽيڪنالاجي, زراعت, صحت, ٻيو) |
check_claim(context, claim, lang="ur"|"sd") |
Urdu, Sindhi | Logical relation (ثابت ٿئي ٿو, متضاد آهي, معلومات ناڪافي آهن) |
consent(text) |
Urdu, Sindhi | Permission state, full authorization |
detect_injection(text) |
Urdu, Sindhi | Text kind, prompt injection attempt |
Ask Custom Questions
You can define your own typed decision questions in any language or script:
from jev_urdu import choice, yes_no
questions = {
"topic": choice(
"هي پيغام ڪهڙي موضوع سان لاڳاپيل آهي؟",
["تعليم", "ٽيڪنالاجي", "زراعت", "صحت", "ٻيو"]
),
"urgent": yes_no("ڇا هي معاملو فوري آهي؟")
}
result = model.ask("اڄ ڪڻڪ جو اگهه وڌي ويو آهي.", questions)
print(result)
Measured Benchmarks & Empirical Findings
Evaluated using controlled, reproducible evaluation suites across Perso-Arabic Sindhi and Roman Sindhi:
1. Zero-Shot Baseline Results:
- Sindhi Script Sentiment: 87.5% (21/24) (
positive: 100%,negative: 87.5%,neutral: 75%). - Sindhi Script Topic Classification: 71.9% (23/32) across 8 domains.
- Sindhi Script Priority / Triage: 75.0% (9/12).
- Sindhi Script Claim / NLI: 66.7% (8/12).
- Roman Sindhi Sentiment (Script-Matched): 66.7% (8/12) (100% precision on neutral and shared cognate roots).
2. Adaptation Improvements:
- Perso-Arabic Sindhi Antonym Contradiction (
گرم ↔ ٿڌي): Jumped from 0% to 100% (2/2) after adaptation. - Roman Sindhi Negative Detection: Jumped from 50% to 75% on authentic vocabulary (
daadho bekaar,mayosi). - Urdu Regression Anchor: 100% preserved (0% regression) across positive, negative, and neutral Urdu controls.
Architecture & Adaptation Details
| Parameter | Specification |
|---|---|
| Encoder | ModernBERT-family bidirectional encoder (jhu-clsp/mmBERT-base, 22 layers, 768 hidden size, 256k-token multilingual vocabulary) |
| Decision Head | 2-layer TransformerEncoderLayer + type embedding + marker-level MLP scorer |
| Parameters | ~322M |
| Supported Scripts | Perso-Arabic Urdu, Roman Urdu, Perso-Arabic Sindhi, Roman Sindhi, Code-mixed English |
| Embeddings Policy | Token embeddings strictly frozen (train_embeddings: false) |
| Fine-tuning Setup | AdamW (encoder_lr: 2.5e-5, head_lr: 1e-4), weight decay 0.01 |
| License | Apache-2.0 |
Lineage and Acknowledgments
- Upstream architecture and base weights: Jev-Urdu by Muhammad Noman / LughaatNLP (Apache-2.0).
- Underlying base encoder: mmBERT-base by JHU CLSP (MIT).
- Sindhi & Roman Sindhi adaptation, presets, and evaluation suite: Developed by Shakeel Ahmed Sanjrani.
- Not affiliated with TypeSafe's commercial Jev API.