YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Romanized Pashto Sentiment Analysis

This model performs sentiment classification on Romanized Pashto text.

Model Description

This model is based on xlm-roberta-base and was fine-tuned for sentiment classification of Romanized Pashto.

Romanized Pashto refers to Pashto written using the Latin/Roman script. Such text can exhibit substantial spelling variation and may contain code-mixing with languages such as Urdu and English, making sentiment classification particularly challenging.

Intended Use

The model is intended for research and experimentation involving Romanized Pashto sentiment analysis and low-resource NLP.

Model Details

  • Base model: xlm-roberta-base
  • Task: Sentiment Classification
  • Language: Pashto
  • Script: Roman/Latin
  • Dataset: Romanized_Pashto_Sentiment
  • Number of classes: 3
  • Classes: Positive, Negative, Neutral

Evaluation

The model achieved a weighted F1 score of:

87.87%

Additional reported results:

  • Weighted F1: 0.878679
  • Training loss: 0.450354
  • Validation loss: 0.357379

Limitations

Romanized Pashto does not have a single standardized spelling system. Variations in spelling, transliteration, code-mixing, and informal writing may affect model performance.

Performance should therefore not be interpreted as representative of all Pashto speakers, domains, or writing styles.

Research Context

This model is part of my broader research on multilingual and low-resource NLP, with a particular interest in Pashto, Romanized and code-mixed language, language resources, and model evaluation.

Downloads last month
27
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support