YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Text Classification Fine-tuning Project

This project fine-tunes a DistilBERT model for sentiment analysis on the IMDB movie review dataset.

What Was Done

  1. Dataset: Used the IMDB dataset from Hugging Face Datasets, which contains 50,000 movie reviews labeled as positive or negative.

  2. Model: Fine-tuned distilbert-base-uncased, a lightweight BERT model, for binary text classification.

  3. Training:

    • Tokenized the dataset with a maximum length of 512 tokens
    • Trained for 1 epoch on a subset of 1,000 training samples
    • Used batch size of 16 with learning rate warmup
    • Evaluated on 200 test samples
  4. Results: The model achieves classification accuracy on the test set. Metrics include accuracy and F1 score.

  5. Deployment: The fine-tuned model is saved locally and can be pushed to the Hugging Face Hub.

Files

  • fine_tune_classifier.py: Main training script
  • push_to_hub.py: Script to push the model to Hugging Face Hub
  • README.md: This file

Usage

Training

pip install transformers datasets torch scikit-learn
python fine_tune_classifier.py

Push to Hugging Face Hub

  1. Install Hugging Face Hub:
pip install huggingface_hub
  1. Login to Hugging Face:
huggingface-cli login
  1. Update the repo_id in push_to_hub.py with your username

  2. Run:

python push_to_hub.py

Using the Model

from transformers import pipeline

classifier = pipeline("text-classification", model="./fine_tuned_model")
result = classifier("This movie was absolutely fantastic!")
print(result)

Model Card

  • Base Model: distilbert-base-uncased
  • Task: Binary Text Classification (Sentiment Analysis)
  • Dataset: IMDB Movie Reviews
  • Labels: 0 (Negative), 1 (Positive)
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support