Transformers from Scratch
Advanced NLP - Assignment 1
HuggingFace repository for Assignment 1 of ANLP 2026 course offered at IIIT Hyderabad in Monsoon 2026.
By: Abdul Kadir Nalawala - 2026701019 - abdulkadir[dot].nalawala[at]research[dot]iiit[dot]ac[dot]in
Training hyper-parameters:
| Hyper-parameter | Value |
|---|---|
seed |
6969 |
batch_size |
32 |
num_epochs |
400 |
lr |
0.0005 |
min_lr_ratio |
0.1 |
warmup_ratio |
0.1 |
weight_decay |
0.01 |
dropout |
0.1 |
early_stopping_patience |
5 |
plateau_ratio |
0.45 |
label_smoothing |
0.2 |
train_ratio |
0.8 |
test_ratio |
0.15 |
validation_ratio |
0.05 |
max_validation_tokens |
512 |
validation_frequency |
200 |
validation_samples |
1 |
tgt_lang |
en |
src_lang |
bin |
seq_len |
512 |
src_vocab_size |
4,096 |
tgt_vocab_size |
1,024 |
tokenization |
bpe |
d_model |
256 |
tokenizer_path |
./tokenizer |
ff |
1,024 |
heads |
4 |
kv_heads |
2 |
layers |
4 |
attention |
MHA/GQA |
Norm |
layernorm/rmsnorm |
pos_encoding |
rope/sinusoidal |
model_filename |
./weights/{config}} |
model_folder |
weights |
Results
| Configuration | Final Training Loss | Final Validation Loss | Epochs Trained |
|---|---|---|---|
| C1 (Base) | 1.287 | 1.013 | 294 |
| C2 (RoPE) | 1.179 | 0.605 | 156 |
| C3 (GQA) | 1.311 | 1.191 | 323 |
| C4 (RMSNorm) | 1.241 | 1.161 | 313 |
| Configuration | Bit Acc. | Seq. Acc. | BLEU (%) | ROUGE-1 | ROUGE-2 | ROUGE-L | ROUGE-Lsum |
|---|---|---|---|---|---|---|---|
| C1 (Base) | 0.658 | 0.005 | 53.172 | 0.754 | 0.605 | 0.748 | 0.748 |
| C2 (RoPE) | 0.704 | 0.023 | 77.086 | 0.885 | 0.820 | 0.882 | 0.881 |
| C3 (GQA) | 0.654 | 0.005 | 49.969 | 0.737 | 0.577 | 0.730 | 0.730 |
| C4 (RMSNorm) | 0.659 | 0.007 | 53.512 | 0.756 | 0.607 | 0.749 | 0.749 |
Repository structure:
ANLP-Assignment1/
βββ tokenizers/
β βββ bin_4096.json ## source tokenizer - binary
β βββ en_1024.json ## target tokenizer - english
β
βββ weights/ # best checkpoints for each config
βββ C1_best.pt
βββ C2_best.pt
βββ C3_best.pt
βββ C4_best.pt
WandB Report
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support