YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
# π¬ Movie Review Sentiment Analyzer
> Fine-tuned FLAN-T5 for binary sentiment classification on movie reviews.





---
## π Overview
This project fine-tunes Google's **FLAN-T5 Base** model on the Stanford IMDB dataset to classify movie reviews as **positive** or **negative**. It includes a full ML pipeline from dataset analysis to a deployed interactive web application.
| | Baseline | Fine-tuned | Improvement |
|---|---|---|---|
| **Accuracy** | 93.50% | **96.00%** | +2.50% β
|
| **F1 Score** | 0.9372 | **0.9596** | +0.0224 β
|
---
## ποΈ Project Structure
sentiment-analysis/ β βββ π 1_dataset_analysis.ipynb # Exploratory data analysis βββ π 2_baseline_evaluation.ipynb # Zero-shot FLAN-T5 evaluation βββ π 3_finetuning.ipynb # Model fine-tuning βββ π 4_posttuning_evaluation.ipynb # Post-tuning evaluation & comparison βββ π 5_app.ipynb # Dashboard HTML generation βββ π server.py # FastAPI inference server βββ π dashboard.html # Interactive web dashboard βββ π baseline_metrics.json # Baseline results βββ π finetuned_metrics.json # Fine-tuned results βββ πΌοΈ *.png # Generated charts
---
## π Quick Start
### 1. Clone the repository
```bash
git clone https://github.com/TerryPotato/clasificador-rese-as-ia.git
cd sentiment-analysis
2. Create and activate virtual environment
python -m venv sentiment_env
sentiment_env\Scripts\activate # Windows
source sentiment_env/bin/activate # Mac/Linux
3. Install dependencies
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128
pip install transformers datasets evaluate accelerate scikit-learn
pip install pandas matplotlib seaborn jupyter
pip install fastapi uvicorn python-multipart
4. Download the fine-tuned model
The model is hosted on HuggingFace Hub. Download it by running this in Python:
from huggingface_hub import snapshot_download
snapshot_download(repo_id="TerryPotato/sentiment-analysis-ai", local_dir="./flan-t5-sentiment-model")
5. Start the server
uvicorn server:app --host 0.0.0.0 --port 8000
6. Open the dashboard
Open dashboard.html in your browser. The green dot confirms the server is online β
π§ Model Details
| Feature | Value |
|---|---|
| Base Model | google/flan-t5-base |
| Parameters | 250M |
| Task | Binary Sentiment Classification |
| Input | Movie review text (English) |
| Output | "positive" or "negative" |
| Max Input Length | 512 tokens |
π§ Fine-tuning Techniques
| Technique | Purpose |
|---|---|
| Early Stopping (patience=2) | Stops training when Validation Loss stops improving, prevents overfitting |
| Cosine LR Scheduler | Gradually reduces learning rate for more stable convergence |
| Gradient Clipping (max_norm=1.0) | Prevents exploding gradients during backpropagation |
| BF16 Precision | Faster training on modern GPUs with numerical stability |
π Training Results
| Epoch | Training Loss | Validation Loss | Status |
|---|---|---|---|
| 1 | 0.1694 | 0.2474 | Decreasing |
| 2 | 0.0632 | 0.1743 | Decreasing |
| 3 | 0.0480 | 0.1450 | β Best Model |
| 4 | 0.0256 | 0.1639 | Overfitting |
| 5 | 0.0261 | 0.1924 | π Early Stop |
π¦ Dataset
| Feature | Value |
|---|---|
| Name | Stanford IMDB Large Movie Review Dataset |
| Source | stanfordnlp/imdb on HuggingFace |
| Total Reviews | 50,000 |
| Class Balance | 50% positive / 50% negative |
| Language | English |
| Used for Fine-tuning | 2,000 reviews (balanced) |
π» Hardware Used
| Component | Spec |
|---|---|
| GPU | NVIDIA RTX 5060 Ti |
| CPU | AMD Ryzen 5 9600X |
| RAM | 32 GB |
| Training Time | ~15 minutes |
π Notebooks Guide
| Notebook | Description |
|---|---|
1_dataset_analysis.ipynb |
Load IMDB dataset, visualize class distribution, word frequency, review lengths |
2_baseline_evaluation.ipynb |
Evaluate FLAN-T5 without fine-tuning on 200 balanced reviews |
3_finetuning.ipynb |
Fine-tune with Early Stopping, Cosine LR, Gradient Clipping |
4_posttuning_evaluation.ipynb |
Compare baseline vs fine-tuned metrics with charts |
5_app.ipynb |
Generate the interactive HTML dashboard |
π₯οΈ Application
The web dashboard includes 4 tabs:
- π Analyze β Submit any review and get real-time sentiment prediction from the model
- π Metrics β Baseline vs Fine-tuned comparison with charts and technique descriptions
- π Training β Epoch-by-epoch loss table and training curve visualization
- π Dataset β IMDB dataset statistics and exploratory analysis charts
π References
- Maas et al. (2011). Learning word vectors for sentiment analysis. ACL.
- Chung et al. (2022). Scaling instruction-finetuned language models. arXiv:2210.11416.
- Wolf et al. (2020). Transformers: State-of-the-art NLP. EMNLP.
---
- Downloads last month
- 3