YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
NEURAL_NETWORK_MODEL_FOR_NEXT_TOKEN_PREDICTION
Abstract
This model is about implementing a Neural Network (FCNN) for next token prediction. The model is trained, validated, and tested using a dataset for the best validation accuracy. The notebook contains code for:
- Data preprocessing
- Model architecture design
- Training the Model
- Evaluating the model performance on test data
Model
The FCNN model is implemented using PyTorch. It uses fully connected layers to learn the mapping between input data and the target labels. The model is loaded from a pre-trained checkpoint (model_small.pth) for evaluation.
Main Components include-
(1) Data Loading: Preprocessed test data (X_test) is fed into the model.
(2) Model Loading: The pre-trained model is loaded using the PyTorch load_state_dict method.
(3) Model Evaluation: The model is evaluated on the test data, and the accuracy is calculated.
The files attached herewith-
Bhavana_FCNN.ipynb: Contains the code for training and evaluating the FCNN model.model_small.pth: Pre-trained model file (used in the evaluation step).
Explanation
Model Evaluation: Load the FCNN model:
load_model = FCNN() load_model.load_state_dict(torch.load('model_small.pth')) load_model.eval()
Pass the test data to the model for predictions:
test_out = load_model(X_test) test_pred = torch.argmax(test_out, dim=1) test_acc = (test_pred == y_test).float().mean() print(f"Test Accuracy: {test_acc.item():.2f}")
Output Expectation:
- The final output will display the test accuracy of the model.
Dependencies
- Python 3.11.5
- PyTorch
- GPU (Optional for faster training)