YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
RuBERT Tiny Spam Classifier π€
A lightweight Russian language spam classifier based on RuBERT Tiny model. The model can detect spam in text messages with high accuracy while maintaining minimal resource requirements.
π Requirements
pip install --upgrade pip
pip install -r requirements.txt
ποΈ Project Structure
.
βββ cleaned_dataset.csv # Cleaned dataset
βββ messages.csv # Source dataset
βββ model/ # Trained model
βββ scripts/
β βββ inference.py # Inference script
β βββ preprocess.ipynb # Data preprocessing notebook
β βββ train.py # Training script
βββ requirements.txt # Project dependencies
Datasets will not be published at this time.
π― Usage
Train
python3 scripts/train.py
Inference
# python3 scripts/inference.py
from scripts.inference import SpamClassifier
classifier = SpamClassifier(model_path="./model")
text = "ΠΡΠΈΠ²Π΅Ρ, Ρ ΠΌΠ΅Π½Ρ Π΅ΡΡΡ ΠΏΠΎΠ΄ΡΠ°Π±ΠΎΡΠΊΠ°, ΠΎΠΏΠ»Π°ΡΠ° 100 ΡΠ΅ΠΊ ΡΡΠ±.
is_spam = classifier.classify(text)
print(f"Is spam: {is_spam}")
π Metrics
- Accuracy: 0.99
- Precision: 0.89
- Recall: 0.96
- F1 Score: 0.92
π Model
The classifier is based on RuBERT Tiny - a lightweight version of RuBERT, optimized for running on low-resource machines. The model is fine-tuned on a dataset of Russian messages for spam classification.
π License
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support