Advanced NLP, Assignment 1

name: Vedant Kulkarni
roll: 2024101140

Links

Setup

Python 3.13 and uv.

uv sync --extra cu126   # or --extra cpu on a machine without a gpu

The two extras select different PyTorch wheels and conflict with each other, so install exactly one.

Download brown_cipher.txt and brown_plain.txt from the dataset link in the assignment and put them in data/. Training reads those two paths directly.

Running

One YAML file under cfg/ describes each configuration. To train a model with a given configuration, run the training script.

uv run python src/train.py cfg/C1/byte-offset.yaml       # base
uv run python src/train.py cfg/C2/byte-offset.yaml       # rope
uv run python src/train.py cfg/C3/byte-offset.yaml       # gqa
uv run python src/train.py cfg/C4/byte-offset.yaml       # rmsnorm
uv run python src/train.py cfg/C5/entropy-patcher.yaml   # blt

Training writes its checkpoint and tokenizers to outputs/<config>/, logs to WandB, and scores its best checkpoint on the test split with greedy decoding when it finishes.

To score a saved checkpoint again, download the weights, then run the evaluation script.

uv run hf download gamemaker1/anlp-assignment-1 --local-dir outputs/
uv run python src/evaluate.py cfg/C5/entropy-patcher.yaml --samples 5

--limit N scores only the first N test lines.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support