YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Natural Language Processing Midterm Project
This README file provides an overview of the content, steps, and code execution for the project described in the accompanying document.
Project Overview
The project demonstrates various machine learning tasks using PyTorch for language processing tasks. The document explains the process of training an RNN model for text generation, evaluation, and result visualization.
Contents of the Project
Data Handling:
- The dataset is read and encoded for model training.
- The dataset includes an English-Basque parallel corpus.
Hyperparameter Setup:
- Hyperparameters include batch size, sequence length, number of iterations, evaluation intervals, learning rate, and embedding sizes.
Model Definition:
- The project defines a simple RNN-based language model using LSTMs and fully connected layers.
- The model is trained to predict the next character in a sequence.
Training and Evaluation:
- The training loop iterates for 200,000 steps, saving checkpoints periodically.
- Both training and validation losses are logged and saved.
Text Generation:
- A function is provided to generate text from the trained model, starting from a prompt (like "Once upon a time").
- Text is generated in both English and Basque languages.
Plotting and Saving Results:
- Training and validation loss curves are plotted and saved as a PNG file.
- Training loss data is saved to a CSV file for further analysis.
Key Steps for Running the Notebook
Import Required Libraries:
- The project relies on PyTorch, Pandas, and Matplotlib.
import torch import pandas as pd import matplotlib.pyplot as pltData Preprocessing:
- The dataset is read and tokenized using character encoding.
with open('/content/eng_basque.txt', 'r', encoding='utf-8') as f: text = f.read() encoded_data = torch.tensor(encode(text), dtype=torch.long)Model Training:
- The SimpleRNNModel is trained using the provided dataset.
model = SimpleRNNModel().to(device) optimizer = torch.optim.AdamW(model.parameters(), lr=learning_rate)Model Evaluation:
- Evaluation functions and loss estimation are defined.
def estimate_loss(): # Function to estimate train/val lossesText Generation:
- After training, text can be generated using the trained model.
start_text = "Once upon a time" generated_text = generate_text(loaded_model, start_text, max_new_tokens=500) print(generated_text)
Generated Results
English Text Example:
Once upon a time...Basque Text Example:
Gaur egun...
Requirements
- Python 3.7+
- PyTorch
- Pandas
- Matplotlib
How to Run the Code
- Clone the project or open the notebook in Google Colab.
- Install the necessary dependencies.
pip install torch pandas matplotlib - Upload the dataset to the notebook and run the cells in sequence.
- After training, generated text and results will be available.
Conclusion
This project provides an end-to-end example of how to build, train, and evaluate a simple RNN for text generation in both English and Basque. The notebook includes code for data preprocessing, model training, checkpoint saving, and result visualization.
Author
- Giri Chandragiri - October 12, 2024