YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Transformers from Scratch & Byte Latent Transformers

Advanced NLP β€” Assignment 1
Transformers from Scratch Β· RoPE Β· GQA Β· RMSNorm Β· BLT


πŸ‘€ Author

Shourya Pillai Β· MS Research, CSE Β· IIIT Hyderabad
Roll No. 2026701047 Β· shourya.pillai@research.iiit.ac.in


πŸš€ Overview

This project explores Transformer architectures for encrypted-text reconstruction, comparing standard subword Transformers against a Byte Latent Transformer (BLT) operating directly on raw bytes.

All Transformer components were implemented from scratch in PyTorch, without using nn.Transformer or nn.MultiheadAttention.

The experiment consists of five controlled configurations:

Model Main Change Representation
C1 Baseline Transformer Custom BPE
C2 RoPE Custom BPE
C3 GQA Custom BPE
C4 RMSNorm Custom BPE
C5 Byte Latent Transformer Raw Bytes

πŸ“Š Results

Final Test Performance

Model Bit Accuracy ↑ Sequence Accuracy ↑ Levenshtein ↓ BLEU ↑
C1 69.67% 1.14% 36.57 0.361
C2 74.48% 3.54% 21.38 0.549
C3 70.63% 1.40% 35.91 0.367
C4 69.44% 1.40% 39.05 0.337
C5 β€” BLT 99.99% 98.60% 0.036 β€”

C5 achieves 99.988% bit-level accuracy and 98.599% exact sequence accuracy.

BLEU and ROUGE are not applicable to C5 because it operates directly on bytes rather than subword tokens.


🧠 What Changed?

C1 β€” Baseline

Sinusoidal absolute positional encoding + Multi-Head Attention + LayerNorm + custom BPE.

C2 β€” RoPE

Replaces absolute positional encoding with Rotary Positional Encoding.

β†’ Best-performing subword Transformer.

C3 β€” GQA

Replaces standard MHA with Grouped-Query Attention, reducing the number of key/value heads while retaining multiple query heads.

C4 β€” RMSNorm

Replaces LayerNorm with RMSNorm while keeping the remaining architecture unchanged.

C5 β€” BLT

Moves from subword tokens to raw bytes, using a Byte Latent Transformer architecture with local byte processing and global Transformer representations.

β†’ Near-perfect reconstruction.


πŸ“ˆ Training

Training Loss

Training Loss

Validation Loss

Validation Loss

Learning Rate

Learning Rate

Peak GPU Memory

Peak GPU Memory


πŸ”‘ Key Takeaways

1. RoPE improves the standard subword Transformer.

C2 substantially outperforms the baseline C1 across bit accuracy, sequence accuracy, Levenshtein distance, BLEU, and ROUGE.

2. GQA is competitive with the baseline but does not improve reconstruction quality in this experiment.

3. RMSNorm does not outperform LayerNorm for this task.

4. Byte-level modeling is extremely effective for this reconstruction problem.

C5 dramatically outperforms all subword-based configurations, achieving near-perfect reconstruction.


πŸ“¦ Repository Contents

C1/                 C1 checkpoint
C2/                 C2 checkpoint
C3/                 C3 checkpoint
C4/                 C4 checkpoint
C5/                 C5 BLT checkpoint

tokenizers/
β”œβ”€β”€ tokenizer.json
└── cipher_tokenizer.json

metrics/
└── evaluation_results.csv

plots/
β”œβ”€β”€ C1_C5_Train_Loss.png
β”œβ”€β”€ C1_C5_Validation_Loss.png
β”œβ”€β”€ C1_C5_Learning_Rate.png
└── Peak_GPU_Memory.png
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support