YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Marathi + English Emotion Classification — IndicBERT

A Marathi + English emotion classification model fine-tuned from ai4bharat/IndicBERTv2-MLM-only.

The model predicts one of seven emotions:

  • Anger
  • Dissatisfaction
  • Happiness
  • Hope
  • Neutral
  • Sadness
  • Satisfaction

Model Details

  • Base Model: ai4bharat/IndicBERTv2-MLM-only
  • Task: Emotion Classification
  • Languages: Marathi and English
  • Number of Classes: 7
  • Maximum Sequence Length: 128
  • Framework: PyTorch
  • Library: Hugging Face Transformers

Emotion Labels

ID Emotion
0 Anger
1 Dissatisfaction
2 Happiness
3 Hope
4 Neutral
5 Sadness
6 Satisfaction

Training Configuration

Parameter Value
Epochs 5
Learning Rate 2e-5
Train Batch Size 16
Evaluation Batch Size 32
Weight Decay 0.01
Warmup Ratio 0.10
Gradient Accumulation 1
Early Stopping Patience 2
Random Seed 42
Maximum Sequence Length 128

Evaluation Results

Metric Score
Accuracy 0.8354
Macro Precision 0.8384
Macro Recall 0.8356
Macro F1 0.8350
Weighted Precision 0.8373
Weighted Recall 0.8354
Weighted F1 0.8343

Per-Class Results

Emotion Precision Recall F1
Anger 0.8805 0.8468 0.8633
Dissatisfaction 0.8154 0.8497 0.8322
Happiness 0.8032 0.8706 0.8356
Hope 0.8840 0.9056 0.8946
Neutral 0.8578 0.6772 0.7569
Sadness 0.8625 0.8776 0.8700
Satisfaction 0.7655 0.8217 0.7926

Usage

Install the required packages:

pip install transformers torch

Load the Model

from transformers import AutoTokenizer, AutoModelForSequenceClassification

model_id = "omgavali26/emotion"

tokenizer = AutoTokenizer.from_pretrained(model_id)

model = AutoModelForSequenceClassification.from_pretrained(model_id)

print("Model loaded successfully!")

Marathi Emotion Prediction

import torch

text = "मला आज खूप आनंद झाला."

inputs = tokenizer(
    text,
    return_tensors="pt",
    truncation=True,
    max_length=128
)

with torch.no_grad():
    outputs = model(**inputs)

predicted_id = torch.argmax(outputs.logits, dim=-1).item()

predicted_label = model.config.id2label[predicted_id]

confidence = torch.softmax(
    outputs.logits,
    dim=-1
)[0][predicted_id].item()

print("Predicted emotion:", predicted_label)
print("Confidence:", confidence)

English Emotion Prediction

import torch

text = "I am very happy today."

inputs = tokenizer(
    text,
    return_tensors="pt",
    truncation=True,
    max_length=128
)

with torch.no_grad():
    outputs = model(**inputs)

predicted_id = torch.argmax(outputs.logits, dim=-1).item()

predicted_label = model.config.id2label[predicted_id]

confidence = torch.softmax(
    outputs.logits,
    dim=-1
)[0][predicted_id].item()

print("Predicted emotion:", predicted_label)
print("Confidence:", confidence)

Using the Transformers Pipeline

from transformers import pipeline

classifier = pipeline(
    "text-classification",
    model="omgavali26/emotion",
    tokenizer="omgavali26/emotion"
)

result = classifier("I am very happy today.")

print(result)

Get Scores for All Emotions

from transformers import pipeline

classifier = pipeline(
    "text-classification",
    model="omgavali26/emotion",
    tokenizer="omgavali26/emotion",
    top_k=7
)

result = classifier("I am very happy today.")

for item in result[0]:
    print(item)

Model Architecture

The model uses:

Input Text
    ↓
IndicBERT Tokenizer
    ↓
IndicBERT Encoder
    ↓
Classification Head
    ↓
7 Emotion Classes

Supported Emotions

The model predicts:

Anger
Dissatisfaction
Happiness
Hope
Neutral
Sadness
Satisfaction

Long Text

The model was trained with a maximum sequence length of 128 tokens.

For longer text, the original inference workflow uses overlapping chunks with:

  • Maximum length: 128
  • Overlap: 32 tokens

Long documents should therefore be divided into smaller chunks before prediction.

Intended Use

This model is intended for:

  • Marathi emotion classification
  • English emotion classification
  • Marathi + English text classification
  • Emotion analysis
  • NLP research
  • Academic projects
  • Sentiment and emotion related applications

Training Environment

The model was fine-tuned using:

  • Python
  • PyTorch
  • Hugging Face Transformers
  • Hugging Face Tokenizers
  • Google Colab

Base Model

This model was fine-tuned from:

ai4bharat/IndicBERTv2-MLM-only

Acknowledgements

Thanks to the AI4Bharat team for the IndicBERT model.

Citation

If you use this model in your project or research, please cite the original IndicBERT model and the dataset used for fine-tuning.

Downloads last month
57
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support