YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Bilingual Language Model (English-Czech)
Project Overview
This project implements a neural network-based language model for next token prediction in both English and Czech. The model uses an LSTM architecture and is trained on a combined dataset of English and Czech text.
Key Features
- Bilingual capability (English and Czech)
- LSTM-based neural network architecture
- SentencePiece tokenization for handling both languages
- Training with perplexity evaluation
Model Architecture
- Embedding layer
- Multi-layer LSTM
- Dropout for regularization
- Fully connected output layer
Dataset
- English: Derived from the Alpaca dataset
- Czech: Custom dataset (details to be specified)
Training Process
- Combined and shuffled English and Czech data
- Trained for 5000 iterations
- Evaluated loss and perplexity at regular intervals
Current Performance
- Final Perplexity: 1900.06
Understanding the Perplexity Score
The perplexity score of 1900.06 indicates that the model is performing reasonably well for a bilingual task, but there's room for improvement. Here's what this score means:
Bilingual Complexity: A perplexity of 1900.06 for a bilingual model is actually quite reasonable. Bilingual models typically have higher perplexity than monolingual models due to the increased complexity of handling two languages simultaneously.
Translation Challenges: The perplexity score suggests that while the model has learned patterns in both languages, it may struggle with precise translations or generating highly coherent text, especially in the less represented language (likely Czech in this case).
Comparison to Monolingual Models: For context, state-of-the-art monolingual models can achieve perplexities below 20, but these are much larger models trained on vast amounts of data.
Implications for Text Generation: With this perplexity, the model can generate text that follows general patterns of both languages but may produce some nonsensical or incorrect phrases, especially when attempting to switch between languages or translate.
Why Translations May Not Be Correct
The perplexity of 89.06 correlates with the observed issues in translation quality:
Vocabulary Limitations: The model may not have a comprehensive grasp of vocabulary in both languages, leading to incorrect word choices.
Contextual Understanding: A higher perplexity indicates that the model sometimes struggles to predict the next token accurately, which can result in contextually inappropriate words or phrases in the generated text.
Grammar and Structure: The model may not have fully captured the grammatical structures of both languages, especially Czech, which has a more complex grammar than English.
Language Mixing: In bilingual settings, the model might inadvertently mix elements from both languages, leading to nonsensical translations.
Data Imbalance: If one language (likely English) was more represented in the training data, the model's performance on the other language (Czech) could be compromised.
Conclusion
While the current model shows promise in bilingual text generation, the perplexity score of 1900.06 indicates that there's significant room for improvement, especially in translation accuracy and coherence. Future iterations of this project should focus on reducing perplexity to enhance the quality of generated text in both languages.