YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Text Classification Fine-tuning Project
This project fine-tunes a DistilBERT model for sentiment analysis on the IMDB movie review dataset.
What Was Done
Dataset: Used the IMDB dataset from Hugging Face Datasets, which contains 50,000 movie reviews labeled as positive or negative.
Model: Fine-tuned
distilbert-base-uncased, a lightweight BERT model, for binary text classification.Training:
- Tokenized the dataset with a maximum length of 512 tokens
- Trained for 1 epoch on a subset of 1,000 training samples
- Used batch size of 16 with learning rate warmup
- Evaluated on 200 test samples
Results: The model achieves classification accuracy on the test set. Metrics include accuracy and F1 score.
Deployment: The fine-tuned model is saved locally and can be pushed to the Hugging Face Hub.
Files
fine_tune_classifier.py: Main training scriptpush_to_hub.py: Script to push the model to Hugging Face HubREADME.md: This file
Usage
Training
pip install transformers datasets torch scikit-learn
python fine_tune_classifier.py
Push to Hugging Face Hub
- Install Hugging Face Hub:
pip install huggingface_hub
- Login to Hugging Face:
huggingface-cli login
Update the
repo_idinpush_to_hub.pywith your usernameRun:
python push_to_hub.py
Using the Model
from transformers import pipeline
classifier = pipeline("text-classification", model="./fine_tuned_model")
result = classifier("This movie was absolutely fantastic!")
print(result)
Model Card
- Base Model: distilbert-base-uncased
- Task: Binary Text Classification (Sentiment Analysis)
- Dataset: IMDB Movie Reviews
- Labels: 0 (Negative), 1 (Positive)
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support