YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Romanized Pashto Sentiment Analysis
This model performs sentiment classification on Romanized Pashto text.
Model Description
This model is based on xlm-roberta-base and was fine-tuned for
sentiment classification of Romanized Pashto.
Romanized Pashto refers to Pashto written using the Latin/Roman script. Such text can exhibit substantial spelling variation and may contain code-mixing with languages such as Urdu and English, making sentiment classification particularly challenging.
Intended Use
The model is intended for research and experimentation involving Romanized Pashto sentiment analysis and low-resource NLP.
Model Details
- Base model: xlm-roberta-base
- Task: Sentiment Classification
- Language: Pashto
- Script: Roman/Latin
- Dataset: Romanized_Pashto_Sentiment
- Number of classes: 3
- Classes: Positive, Negative, Neutral
Evaluation
The model achieved a weighted F1 score of:
87.87%
Additional reported results:
- Weighted F1: 0.878679
- Training loss: 0.450354
- Validation loss: 0.357379
Limitations
Romanized Pashto does not have a single standardized spelling system. Variations in spelling, transliteration, code-mixing, and informal writing may affect model performance.
Performance should therefore not be interpreted as representative of all Pashto speakers, domains, or writing styles.
Research Context
This model is part of my broader research on multilingual and low-resource NLP, with a particular interest in Pashto, Romanized and code-mixed language, language resources, and model evaluation.
- Downloads last month
- 27