YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Natural Language Processing Midterm Project

This README file provides an overview of the content, steps, and code execution for the project described in the accompanying document.

Project Overview

The project demonstrates various machine learning tasks using PyTorch for language processing tasks. The document explains the process of training an RNN model for text generation, evaluation, and result visualization.


Contents of the Project

  1. Data Handling:

    • The dataset is read and encoded for model training.
    • The dataset includes an English-Basque parallel corpus.
  2. Hyperparameter Setup:

    • Hyperparameters include batch size, sequence length, number of iterations, evaluation intervals, learning rate, and embedding sizes.
  3. Model Definition:

    • The project defines a simple RNN-based language model using LSTMs and fully connected layers.
    • The model is trained to predict the next character in a sequence.
  4. Training and Evaluation:

    • The training loop iterates for 200,000 steps, saving checkpoints periodically.
    • Both training and validation losses are logged and saved.
  5. Text Generation:

    • A function is provided to generate text from the trained model, starting from a prompt (like "Once upon a time").
    • Text is generated in both English and Basque languages.
  6. Plotting and Saving Results:

    • Training and validation loss curves are plotted and saved as a PNG file.
    • Training loss data is saved to a CSV file for further analysis.

Key Steps for Running the Notebook

  1. Import Required Libraries:

    • The project relies on PyTorch, Pandas, and Matplotlib.
    import torch
    import pandas as pd
    import matplotlib.pyplot as plt
    
  2. Data Preprocessing:

    • The dataset is read and tokenized using character encoding.
    with open('/content/eng_basque.txt', 'r', encoding='utf-8') as f:
        text = f.read()
    
    encoded_data = torch.tensor(encode(text), dtype=torch.long)
    
  3. Model Training:

    • The SimpleRNNModel is trained using the provided dataset.
    model = SimpleRNNModel().to(device)
    optimizer = torch.optim.AdamW(model.parameters(), lr=learning_rate)
    
  4. Model Evaluation:

    • Evaluation functions and loss estimation are defined.
    def estimate_loss():
        # Function to estimate train/val losses
    
  5. Text Generation:

    • After training, text can be generated using the trained model.
    start_text = "Once upon a time"
    generated_text = generate_text(loaded_model, start_text, max_new_tokens=500)
    print(generated_text)
    

Generated Results

  • English Text Example:

    Once upon a time...
    
  • Basque Text Example:

    Gaur egun...
    

Requirements

  • Python 3.7+
  • PyTorch
  • Pandas
  • Matplotlib

How to Run the Code

  1. Clone the project or open the notebook in Google Colab.
  2. Install the necessary dependencies.
    pip install torch pandas matplotlib
    
  3. Upload the dataset to the notebook and run the cells in sequence.
  4. After training, generated text and results will be available.

Conclusion

This project provides an end-to-end example of how to build, train, and evaluate a simple RNN for text generation in both English and Basque. The notebook includes code for data preprocessing, model training, checkpoint saving, and result visualization.


Author

  • Giri Chandragiri - October 12, 2024
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support