YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
This project applies BERT (Bidirectional Encoder Representations from Transformers) to classify movie reviews into Positive or Negative sentiment. It includes:
Data loading and EDA
Text preprocessing
BERT tokenization
Custom PyTorch dataset
Fine-tuning a pre-trained BERT model
Evaluation (Accuracy, F1-score, Confusion Matrix)
Plots saved in /plots folder
This is my first advanced NLP project, and it represents a key step in my growth from traditional Machine Learning to modern Deep Learning techniques used in real companies.
📂 Project Structure Proyecto6_SentimentBERT/ │ ├── reviews.csv # Dataset (movie reviews + sentiment) ├── notebook.ipynb # Full training and evaluation pipeline ├── plots/ │ ├── sentiment_distribution.png │ ├── text_length_distribution.png │ └── bert_confusion_matrix.png └── README.md
🧠 Dataset
The dataset contains 1000 movie reviews:
500 positive (1)
500 negative (0)
Each entry contains:
review: text content
sentiment: label
🔍 Exploratory Data Analysis (EDA)
The notebook includes:
✔ Class distribution
Balanced dataset (500/500).
✔ Text length analysis
Helps adjust tokenization settings.
Plots are saved automatically in the plots/ directory.
🤖 BERT Model
Model used:
bert-base-uncased
Steps:
Tokenization with BertTokenizer
Custom PyTorch Dataset
Train/test split
Fine-tuning BERT using Trainer
Evaluation on test set
📊 Results ✔ Test Accuracy: 100% ✔ F1-score: 1.0 Confusion Matrix
Perfect classification:
True \ Pred 0 1 0 (neg) 100 0 1 (pos) 0 100
Saved as plots/bert_confusion_matrix.png.
🛠️ Technologies Used
Python
PyTorch
Transformers (HuggingFace)
NumPy
Pandas
Matplotlib
Seaborn
🚀 How to Run
Install requirements:
pip install transformers torch accelerate
Open the notebook:
jupyter notebook SENTIMENTBERT.ipynb
Run cells in order. (Training on CPU works, but GPU is faster.)
(README en español) Este proyecto utiliza BERT, uno de los modelos más potentes de NLP, para clasificar reseñas de películas como Positivas o Negativas.
Incluye:
Carga del dataset
EDA
Preprocesado de texto
Tokenización con BERT
Creación de un Dataset personalizado
Fine-tuning de BERT
Evaluación con métricas avanzadas
Gráficos guardados en /plots
Es mi primer proyecto avanzado con NLP y Deep Learning, un paso importante más allá del Machine Learning tradicional.
📊 Resultados ✔ Precisión en test: 100% ✔ F1-score: 1.0 Matriz de confusión
Clasificación perfecta de las 200 muestras del test.
🛠️ Tecnologías
Python
PyTorch
Transformers (HuggingFace)
Pandas
NumPy
Matplotlib / Seaborn
📥 Ejecución
Instalar dependencias:
pip install transformers torch accelerate
Ejecutar el notebook:
jupyter notebook SENTIMENTBERT.ipynb
- Downloads last month
- 5