Advanced NLP, Assignment 1
name: Vedant Kulkarni
roll: 2024101140
Links
Setup
Python 3.13 and uv.
uv sync --extra cu126 # or --extra cpu on a machine without a gpu
The two extras select different PyTorch wheels and conflict with each other, so install exactly one.
Download brown_cipher.txt and brown_plain.txt from the dataset link
in the assignment and put them in data/. Training reads those two
paths directly.
Running
One YAML file under cfg/ describes each configuration. To train a
model with a given configuration, run the training script.
uv run python src/train.py cfg/C1/byte-offset.yaml # base
uv run python src/train.py cfg/C2/byte-offset.yaml # rope
uv run python src/train.py cfg/C3/byte-offset.yaml # gqa
uv run python src/train.py cfg/C4/byte-offset.yaml # rmsnorm
uv run python src/train.py cfg/C5/entropy-patcher.yaml # blt
Training writes its checkpoint and tokenizers to outputs/<config>/,
logs to WandB, and scores its best checkpoint on the test split with
greedy decoding when it finishes.
To score a saved checkpoint again, download the weights, then run the evaluation script.
uv run hf download gamemaker1/anlp-assignment-1 --local-dir outputs/
uv run python src/evaluate.py cfg/C5/entropy-patcher.yaml --samples 5
--limit N scores only the first N test lines.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support