Jev-Multilingual - Urdu and Sindhi Decisions

Jev-Multilingual: Urdu, Sindhi & Roman Decisions for Python

A multilingual extension of muhammadnoman76/jev-urdu, adding native Sindhi (سنڌي) and Roman Sindhi decision-making capabilities to the Jev architecture.

Base model developed by Muhammad Noman through LughaatNLP.
Sindhi and Roman Sindhi adaptation, presets, and benchmarks developed by Shakeel Ahmed Sanjrani.

Model Fork | Upstream Base Model | Apache-2.0 license | Python 3.10+


From a message to a structured decision

Urdu and Sindhi in customer applications arrive in multiple scripts: Arabic script, Roman script (Latin phonetics), and English mixed text.

jev-multilingual provides a unified, typed decision model across both languages:

import jev_urdu

model = jev_urdu.load("shakeel143/jev-multilingual")

# 1. Native Sindhi Script Triage & Sentiment
sindhi_msg = "منهنجو يونيورسٽي پورٽل ٽن ڏينهن کان لاگ ان نٿو ٿئي، فوري مدد گهرجي."
print(model.triage(sindhi_msg, lang="sd"))
print(model.sentiment("هي سروس تمام بهترين ۽ سٺي آهي.", lang="sd"))

# 2. Native Sindhi Topic Classification
topic_msg = "اسان جي زمين ۾ ڪڻڪ جو فصل ڏاڍو سٺو ٿيو آهي."
print(model.topic(topic_msg, lang="sd"))

# 3. Roman Sindhi Decisions
print(model.sentiment("hee phone daadho sutho aa, battery zabardast aa.", lang="sd"))

# 4. Urdu (100% Backwards Compatible)
print(model.sentiment("یہ فون بہت اچھا ہے، بیٹری بھی زبردست ہے۔", lang="ur"))

The model returns structured decisions and calibrated probability distributions rather than conversational text generations.


What is new in Jev-Multilingual?

  1. Native Sindhi Task Presets (SINDHI_PRESETS):
    • Pre-configured, natural Sindhi question prompts and label spaces for triage, sentiment, claim verification, and topic routing.
  2. Language-Aware API:
    • Built-in sentiment(), triage(), check_claim(), and topic() methods support lang="sd" for Sindhi and lang="ur" for Urdu.
  3. Roman Sindhi Evaluation & Lexical Grounding:
    • Evaluated and adapted on authentic Roman Sindhi vocabulary (daadho, sutho, ghandi, mayosi, bekaar).
  4. Antonym Contradiction Resolution:
    • Fine-tuned decision head resolves zero-shot antonym confusion in Sindhi (e.g. گرم ↔ ٿڌي).
  5. Zero Catastrophic Forgetting:
    • Base Urdu benchmarks remain 100% shielded and regression-tested.

Built-in Tasks

Method Languages Decisions Returned
triage(text, lang="ur"|"sd") Urdu, Sindhi Issue (بنيادي مسئلو), Human request (انساني مدد), Priority (ترجيح)
sentiment(text, lang="ur"|"sd") Urdu, Sindhi, Roman Sentiment (مثبت/منفي/غير جانبدار), Dissatisfaction
topic(text, lang="sd") Sindhi, Roman Sindhi Domain category (تعليم, ٽيڪنالاجي, زراعت, صحت, ٻيو)
check_claim(context, claim, lang="ur"|"sd") Urdu, Sindhi Logical relation (ثابت ٿئي ٿو, متضاد آهي, معلومات ناڪافي آهن)
consent(text) Urdu, Sindhi Permission state, full authorization
detect_injection(text) Urdu, Sindhi Text kind, prompt injection attempt

Ask Custom Questions

You can define your own typed decision questions in any language or script:

from jev_urdu import choice, yes_no

questions = {
    "topic": choice(
        "هي پيغام ڪهڙي موضوع سان لاڳاپيل آهي؟",
        ["تعليم", "ٽيڪنالاجي", "زراعت", "صحت", "ٻيو"]
    ),
    "urgent": yes_no("ڇا هي معاملو فوري آهي؟")
}

result = model.ask("اڄ ڪڻڪ جو اگهه وڌي ويو آهي.", questions)
print(result)

Measured Benchmarks & Empirical Findings

Evaluated using controlled, reproducible evaluation suites across Perso-Arabic Sindhi and Roman Sindhi:

1. Zero-Shot Baseline Results:

  • Sindhi Script Sentiment: 87.5% (21/24) (positive: 100%, negative: 87.5%, neutral: 75%).
  • Sindhi Script Topic Classification: 71.9% (23/32) across 8 domains.
  • Sindhi Script Priority / Triage: 75.0% (9/12).
  • Sindhi Script Claim / NLI: 66.7% (8/12).
  • Roman Sindhi Sentiment (Script-Matched): 66.7% (8/12) (100% precision on neutral and shared cognate roots).

2. Adaptation Improvements:

  • Perso-Arabic Sindhi Antonym Contradiction (گرم ↔ ٿڌي): Jumped from 0% to 100% (2/2) after adaptation.
  • Roman Sindhi Negative Detection: Jumped from 50% to 75% on authentic vocabulary (daadho bekaar, mayosi).
  • Urdu Regression Anchor: 100% preserved (0% regression) across positive, negative, and neutral Urdu controls.

Architecture & Adaptation Details

Parameter Specification
Encoder ModernBERT-family bidirectional encoder (jhu-clsp/mmBERT-base, 22 layers, 768 hidden size, 256k-token multilingual vocabulary)
Decision Head 2-layer TransformerEncoderLayer + type embedding + marker-level MLP scorer
Parameters ~322M
Supported Scripts Perso-Arabic Urdu, Roman Urdu, Perso-Arabic Sindhi, Roman Sindhi, Code-mixed English
Embeddings Policy Token embeddings strictly frozen (train_embeddings: false)
Fine-tuning Setup AdamW (encoder_lr: 2.5e-5, head_lr: 1e-4), weight decay 0.01
License Apache-2.0

Lineage and Acknowledgments

  • Upstream architecture and base weights: Jev-Urdu by Muhammad Noman / LughaatNLP (Apache-2.0).
  • Underlying base encoder: mmBERT-base by JHU CLSP (MIT).
  • Sindhi & Roman Sindhi adaptation, presets, and evaluation suite: Developed by Shakeel Ahmed Sanjrani.
  • Not affiliated with TypeSafe's commercial Jev API.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for shakeel143/jev-multilingual

Finetuned
(1)
this model