ANLP Assignment 2 - Monsoon 2026
Author: Krithik Kambhampati (krithik3001)
Institution: IIIT Hyderabad
Checkpoints, tokenizer, plots and evaluation results for Assignment 2. All numbers below are generated from the result JSON files in this repository.
WandB Projects
- Part 1: https://wandb.ai/krithik-kambhampati-iiit-hyderabad/anlp-assignment2-part1
- Part 2: https://wandb.ai/krithik-kambhampati-iiit-hyderabad/anlp-assignment2-part2
- Part 3: https://wandb.ai/krithik-kambhampati-iiit-hyderabad/anlp-assignment2-part3
Part 1: Mixture of Experts (test split of belumind/en-vi-ja-curated-500k-triplets)
| Variant | Total Params | Active Params | Test Loss | Test PPL | Test BLEU | Checkpoint |
|---|---|---|---|---|---|---|
| dense | 27,422,720 | 27,422,720 | 2.3242 | 10.22 | 24.57 | part1/best_dense.pt |
| moe_top1 | 27,444,224 | 17,988,608 | 2.3633 | 10.63 | 24.35 | part1/best_moe_top1.pt |
| moe_top2 | 27,444,224 | 21,140,480 | 2.3427 | 10.41 | 24.50 | part1/best_moe_top2.pt |
| moe_shared | 27,441,152 | 21,137,408 | 2.2974 | 9.95 | 25.01 | part1/best_moe_shared.pt |
| moe_active_match | 40,039,424 | 27,438,080 | 2.3122 | 10.10 | 24.42 | part1/best_moe_active_match.pt |
Routing heatmaps: assets/routing_heatmaps/.
Part 2: Custom Optimizers (browndw/human-ai-parallel-corpus, final 1.0x milestone)
| Optimizer | LR | Tokens | Val Loss | Val PPL | Continuation BLEU | Checkpoint |
|---|---|---|---|---|---|---|
| adamw | 0.0003 | 52,152,638 | 3.5396 | 34.45 | 1.12 | part2/adamw_1.0x.pt |
| nadamw | 0.0003 | 52,152,638 | 3.5231 | 33.89 | 1.15 | part2/nadamw_1.0x.pt |
| lion | 0.0003 | 52,152,638 | 3.7967 | 44.56 | 0.55 | part2/lion_1.0x.pt |
| muon | 0.02 | 52,152,638 | 3.3953 | 29.82 | 1.20 | part2/muon_1.0x.pt |
| sophia | 0.0003 | 52,152,638 | 4.4071 | 82.03 | 0.47 | part2/sophia_1.0x.pt |
Training dynamics: assets/part2_dynamics/.
Part 3: Decoding Strategies (EleutherAI/Pythia-160M on hamishivi/ROCStories)
Reference (human continuation) perplexity: 25.03. Greedy == beam(W=1) on 100.0% of samples.
| Strategy | PPL | Token Acc (%) | F1 (%) | BLEU | Avg Gen Tokens | ms / sample | ms / token |
|---|---|---|---|---|---|---|---|
| greedy | 2.18 | 1.43 | 13.27 | 0.44 | 64.0 | 274.1 | 4.28 |
| top_k_10 | 7.70 | 1.48 | 19.60 | 0.50 | 63.6 | 264.8 | 4.16 |
| top_k_50 | 13.77 | 1.32 | 18.74 | 0.34 | 63.2 | 246.6 | 3.90 |
| top_p_0.7 | 10.30 | 1.38 | 18.87 | 0.32 | 63.3 | 243.9 | 3.85 |
| top_p_0.9 | 24.45 | 1.17 | 17.42 | 0.30 | 62.9 | 243.2 | 3.87 |
| top_p_0.95 | 34.56 | 1.11 | 16.74 | 0.20 | 63.2 | 244.6 | 3.87 |
| beam_width_1 | 2.18 | 1.43 | 13.27 | 0.44 | 64.0 | 245.9 | 3.84 |
| beam_width_2 | 1.94 | 1.22 | 11.50 | 0.30 | 63.9 | 250.9 | 3.93 |
| beam_width_4 | 1.80 | 1.17 | 10.79 | 0.29 | 63.6 | 256.8 | 4.04 |
Raw generations: assets/part3/rocstories_generations.json.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support