BanglaBERT DAPT — Bangla Depression Classifier (Sub-component of Multimodal System)

Project: Detecting Mental Health and Suicidal Tendencies on Social Media using Multimodal NLP
Institution: BRAC University, Department of CSE
Authors: Md. Nazim Hossain, Kazi Tarif Rahman, Raiseen Jahan Ritu, Ahmad Sameer


Model Description

This is the depression severity sub-classifier that feeds into the multimodal fuzzy-logic risk combination system. It reads Bangla text and outputs a 4-class depression severity distribution, which is then combined with the suicide-severity model's output using fuzzy logic rules to produce a final risk level (Minimal / Low / Elevated / Critical).

Architecture: BanglaBERT (ELECTRA discriminator, 12L, 768H) fine-tuned on 4,897 Bangla depression severity posts after Domain-Adaptive Pretraining (DAPT) on 250,154 translated mental health Reddit posts.


Performance (In-domain evaluation on the Bangla depression dataset)

Metric Value
Accuracy 93.20%
Macro-F1 0.9182
Weighted-F1 0.9308

Per-class F1:

Label Category F1
1 Minimum (Mild/Non-clinical) 0.963
2 Mild (Moderate Depression) 0.924
3 Moderate (Severe/Passive Ideation) 0.832
4 Severe (Active Suicidal Ideation) 0.954

Important caveat: This metric was evaluated on a test split of the same dataset the model was trained on (~80% overlap with training domain). It is a behaviour sanity check confirming the model functions correctly — not an unbiased held-out generalisation estimate. This is explicitly documented in the thesis.


Role in the Multimodal System

Bangla text (meme OCR or raw text)
    → This model (depression classifier)
    → 4-class probability distribution [Minimum, Mild, Moderate, Severe]
         ↓
    Combined with suicide-severity model output
         ↓
    Fuzzy-logic risk layer (AND/OR rules)
         ↓
    Final risk level: Minimal / Low / Elevated / Critical

The model's full probability distribution (not just the top class) is used as input to the fuzzy-logic rules, allowing borderline cases to surface as partial risk signals rather than being clipped to a single hard label.


Loading Instructions

Critical: Load the tokenizer from this repo directly. The checkpoint itself does not contain a fully saved tokenizer — the tokenizer files have been merged from the TAPT checkpoint used in the pipeline.

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

tokenizer = AutoTokenizer.from_pretrained("SrothJr/banglabert-dapt-depression-classifier")
model = AutoModelForSequenceClassification.from_pretrained("SrothJr/banglabert-dapt-depression-classifier")
model.eval()

LABEL_MAP = {0: "Minimum", 1: "Mild", 2: "Moderate", 3: "Severe"}

text = "আমি খুব একা অনুভব করছি।"
inputs = tokenizer(text, return_tensors="pt", max_length=256, truncation=True, padding=True)

with torch.no_grad():
    logits = model(**inputs).logits
    probs = torch.softmax(logits, dim=-1)

predicted_idx = probs.argmax(dim=-1).item()
print(f"Predicted class: {LABEL_MAP[predicted_idx]}")
print(f"Probability distribution: {probs.squeeze().tolist()}")

Training Details

Property Value
Base model csebuetnlp/banglabert
DAPT corpus 250,154 lines, translated from 15 English Reddit mental health subreddits via NLLB-200-3.3B
TAPT Per-fold task-adaptive pretraining on training partition texts (MLM)
Fine-tuning data 4,897 annotated Bangla depression severity posts (Kabir et al.)
Evaluation Stratified 5-fold cross-validation, fold 5 checkpoint reported
Max sequence length 256
Learning rate 2e-5
Effective batch size 16

Limitations

  • The 93.2% metric overlaps its training domain — treat as a sanity check, not a generalisation number.
  • Performance on meme text (short, colloquial, OCR-extracted) has not been independently benchmarked.
  • Not a clinical tool — outputs describe textual content, never a judgment about a real person.

Citation

Part of the thesis: Detecting Mental Health and Suicidal Tendencies on Social Media using Multimodal NLP
BRAC University, Department of CSE, 2026.

Downloads last month
5
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support