YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

This project applies BERT (Bidirectional Encoder Representations from Transformers) to classify movie reviews into Positive or Negative sentiment. It includes:

Data loading and EDA

Text preprocessing

BERT tokenization

Custom PyTorch dataset

Fine-tuning a pre-trained BERT model

Evaluation (Accuracy, F1-score, Confusion Matrix)

Plots saved in /plots folder

This is my first advanced NLP project, and it represents a key step in my growth from traditional Machine Learning to modern Deep Learning techniques used in real companies.

📂 Project Structure Proyecto6_SentimentBERT/ │ ├── reviews.csv # Dataset (movie reviews + sentiment) ├── notebook.ipynb # Full training and evaluation pipeline ├── plots/ │ ├── sentiment_distribution.png │ ├── text_length_distribution.png │ └── bert_confusion_matrix.png └── README.md

🧠 Dataset

The dataset contains 1000 movie reviews:

500 positive (1)

500 negative (0)

Each entry contains:

review: text content

sentiment: label

🔍 Exploratory Data Analysis (EDA)

The notebook includes:

✔ Class distribution

Balanced dataset (500/500).

✔ Text length analysis

Helps adjust tokenization settings.

Plots are saved automatically in the plots/ directory.

🤖 BERT Model

Model used:

bert-base-uncased

Steps:

Tokenization with BertTokenizer

Custom PyTorch Dataset

Train/test split

Fine-tuning BERT using Trainer

Evaluation on test set

📊 Results ✔ Test Accuracy: 100% ✔ F1-score: 1.0 Confusion Matrix

Perfect classification:

True \ Pred 0 1 0 (neg) 100 0 1 (pos) 0 100

Saved as plots/bert_confusion_matrix.png.

🛠️ Technologies Used

Python

PyTorch

Transformers (HuggingFace)

NumPy

Pandas

Matplotlib

Seaborn

🚀 How to Run

Install requirements:

pip install transformers torch accelerate

Open the notebook:

jupyter notebook SENTIMENTBERT.ipynb

Run cells in order. (Training on CPU works, but GPU is faster.)

(README en español) Este proyecto utiliza BERT, uno de los modelos más potentes de NLP, para clasificar reseñas de películas como Positivas o Negativas.

Incluye:

Carga del dataset

EDA

Preprocesado de texto

Tokenización con BERT

Creación de un Dataset personalizado

Fine-tuning de BERT

Evaluación con métricas avanzadas

Gráficos guardados en /plots

Es mi primer proyecto avanzado con NLP y Deep Learning, un paso importante más allá del Machine Learning tradicional.

📊 Resultados ✔ Precisión en test: 100% ✔ F1-score: 1.0 Matriz de confusión

Clasificación perfecta de las 200 muestras del test.

🛠️ Tecnologías

Python

PyTorch

Transformers (HuggingFace)

Pandas

NumPy

Matplotlib / Seaborn

📥 Ejecución

Instalar dependencias:

pip install transformers torch accelerate

Ejecutar el notebook:

jupyter notebook SENTIMENTBERT.ipynb

Downloads last month
5
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support