mohjkhan/anlp-a2
Checkpoints for ANLP Assignment 2 (Mixture of Experts, Optimizers, Decoding Strategies). Each checkpoint is a plain torch.save dict with keys model (state_dict) and config (the TransformerConfig). Load with the training repo's Transformer class; no transformers/AutoModel wrapper is provided.
Part 1 β FFN / MoE variants
Decoder-only transformer (dim 512, 8 layers, 8 heads, SwiGLU, 32k vocab, ctx 512, RoPE) trained on belumind/en-vi-ja-curated-500k-triplets (VI+JAβEN) at a matched 30M-token budget.
Part 2 β Optimizers
Same dense model pretrained on browndw/human-ai-parallel-corpus to ~1Γ Chinchilla (β38.5M tokens). Final values:
| Optimizer | Val loss | Val ppl | Test BLEU |
|---|---|---|---|
| adamw | 4.53 | 92.7 | 0.703 |
| mars | 4.62 | 101.1 | 0.768 |
| adafactor | 5.80 | 331.0 | 0.487 |
| muon | 4.62 | 101.9 | 0.774 |
| sophia | 4.94 | 140.0 | 0.532 |
Part 3 β Decoding
Decoding strategies (greedy, top-k, top-p, beam 1/2/4) were evaluated on EleutherAI/pythia-160m with hamishivi/ROCStories and do not produce trained checkpoints (see the code repo + W&B for results).