YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Transformers from Scratch & Byte Latent Transformers
Advanced NLP β Assignment 1
Transformers from Scratch Β· RoPE Β· GQA Β· RMSNorm Β· BLT
π€ Author
Shourya Pillai Β· MS Research, CSE Β· IIIT Hyderabad
Roll No. 2026701047 Β· shourya.pillai@research.iiit.ac.in
π Overview
This project explores Transformer architectures for encrypted-text reconstruction, comparing standard subword Transformers against a Byte Latent Transformer (BLT) operating directly on raw bytes.
All Transformer components were implemented from scratch in PyTorch,
without using nn.Transformer or nn.MultiheadAttention.
The experiment consists of five controlled configurations:
| Model | Main Change | Representation |
|---|---|---|
| C1 | Baseline Transformer | Custom BPE |
| C2 | RoPE | Custom BPE |
| C3 | GQA | Custom BPE |
| C4 | RMSNorm | Custom BPE |
| C5 | Byte Latent Transformer | Raw Bytes |
π Results
Final Test Performance
| Model | Bit Accuracy β | Sequence Accuracy β | Levenshtein β | BLEU β |
|---|---|---|---|---|
| C1 | 69.67% | 1.14% | 36.57 | 0.361 |
| C2 | 74.48% | 3.54% | 21.38 | 0.549 |
| C3 | 70.63% | 1.40% | 35.91 | 0.367 |
| C4 | 69.44% | 1.40% | 39.05 | 0.337 |
| C5 β BLT | 99.99% | 98.60% | 0.036 | β |
C5 achieves 99.988% bit-level accuracy and 98.599% exact sequence accuracy.
BLEU and ROUGE are not applicable to C5 because it operates directly on bytes rather than subword tokens.
π§ What Changed?
C1 β Baseline
Sinusoidal absolute positional encoding + Multi-Head Attention + LayerNorm + custom BPE.
C2 β RoPE
Replaces absolute positional encoding with Rotary Positional Encoding.
β Best-performing subword Transformer.
C3 β GQA
Replaces standard MHA with Grouped-Query Attention, reducing the number of key/value heads while retaining multiple query heads.
C4 β RMSNorm
Replaces LayerNorm with RMSNorm while keeping the remaining architecture unchanged.
C5 β BLT
Moves from subword tokens to raw bytes, using a Byte Latent Transformer architecture with local byte processing and global Transformer representations.
β Near-perfect reconstruction.
π Training
Training Loss
Validation Loss
Learning Rate
Peak GPU Memory
π Key Takeaways
1. RoPE improves the standard subword Transformer.
C2 substantially outperforms the baseline C1 across bit accuracy, sequence accuracy, Levenshtein distance, BLEU, and ROUGE.
2. GQA is competitive with the baseline but does not improve reconstruction quality in this experiment.
3. RMSNorm does not outperform LayerNorm for this task.
4. Byte-level modeling is extremely effective for this reconstruction problem.
C5 dramatically outperforms all subword-based configurations, achieving near-perfect reconstruction.
π¦ Repository Contents
C1/ C1 checkpoint
C2/ C2 checkpoint
C3/ C3 checkpoint
C4/ C4 checkpoint
C5/ C5 BLT checkpoint
tokenizers/
βββ tokenizer.json
βββ cipher_tokenizer.json
metrics/
βββ evaluation_results.csv
plots/
βββ C1_C5_Train_Loss.png
βββ C1_C5_Validation_Loss.png
βββ C1_C5_Learning_Rate.png
βββ Peak_GPU_Memory.png



