ANLP Assignment 2: Mixture of Experts (MoE), Pretraining Optimizers & Decoding Strategies

This repository hosts the trained model checkpoints, custom tokenizer, configurations, and evaluation outputs for ANLP Assignment 2 (Monsoon 2026).

πŸ“Œ Repository Overview

  • Part 1 β€” Mixture of Experts (MoE): 5 ablation variants of a decoder-only transformer with dense MLP vs. Top-1, Top-2, Shared+Routed, and parameter-matched MoE FFN layers trained on belumind/en-vi-ja-curated-500k-triplets for 90M tokens (3x baseline budget).
  • Part 2 β€” Pretraining Optimizers: Comprehensive empirical benchmarking of 5 distinct optimizer families (AdamW, Cautious AdamW, Lion, Muon, and Sophia-G) trained on browndw/human-ai-parallel-corpus for 30M tokens.
  • Part 3 β€” Decoding Strategies: Benchmark of Greedy, Beam Search ($W \in {1, 2, 4, 8}$), Top-$k$ sampling ($k=20, 50$), and Top-$p$ (Nucleus) sampling ($p=0.80, 0.90, 0.95$) evaluating repetition rate, distinct $n$-grams, BLEU/ROUGE overlap, and latency.

πŸ“Š Part 1: Mixture-of-Experts Benchmark Results

All variants share an identical backbone ($D_{model}=512$, $N_{layers}=6$, $N_{heads}=8$, RMSNorm, RoPE, Vocab=16,384) and were trained on an identical 90,000,000 token budget:

Variant Description Total Params Active Params Val Loss Val PPL Test BLEU Routing Strategy
V1 Standard Dense 2-Layer MLP 35.69M 35.67M 3.3361 28.11 0.6585 Dense Baseline ($d_{ff}=2048$)
V2 MoE (4 experts, Top-1 active) 54.58M 35.67M 3.3980 29.90 0.6653 Top-1 softmax routing
V3 MoE (4 experts, Top-2 active) 54.58M 41.97M 3.0185 20.46 0.7188 Top-2 normalized routing
V4 Shared + Routed (1 shared + Top-1/3 routed) 54.58M 41.97M 3.0336 20.77 0.7099 1 always-active + Top-1 routed
V5 Parameter-Matched MoE (4 experts, Top-2) 45.13M 35.67M 3.2085 24.74 0.6865 Top-2 ($d_{ff}=1024$, matched active FLOPs)

⚑ Part 2: Pretraining Optimizers Benchmark Results

Comparison of all 5 optimizers on browndw/human-ai-parallel-corpus under an identical 30M token budget:

Category Optimizer Best LR State Memory Final Val Loss Final Val PPL Final BLEU Speedup to 3.5 Loss
Cat 1: Standard Adaptive AdamW $1 \times 10^{-3}$ $2\times$ (8 B/param) 3.4309 30.90 0.6133 $1.00\times$ (0.9x tokens)
Cat 2: Variance-Reduced Cautious AdamW $1 \times 10^{-3}$ $2\times$ (8 B/param) 3.4209 30.60 0.7022 $1.05\times$ (0.8x tokens)
Cat 3: Memory-Efficient Lion $3 \times 10^{-4}$ $1\times$ (4 B/param) 3.6085 36.91 0.6933 $0.90\times$ (1.0x tokens)
Cat 4: Matrix-Based Muon $1 \times 10^{-2}$ $1\times$ (4 B/param) 3.3077 27.32 0.7297 $\mathbf{1.35\times}$ (0.7x tokens)
Cat 5: Hessian-Based (Bonus) Sophia-G $5 \times 10^{-4}$ $2\times$ (8 B/param) 4.4871 88.86 0.3965 $> 1.0\times$ (did not reach)

πŸ› οΈ How to Load and Inspect Checkpoints

1. Download Model Weights via huggingface_hub

from huggingface_hub import hf_hub_download
import torch

# Download Variant 3 best checkpoint
ckpt_path = hf_hub_download(
    repo_id="anuml/anlp-assignment2-models",
    filename="checkpoints/part1/variant3/model_variant3_best.pt"
)

# Load checkpoint state dict & training metadata
checkpoint = torch.load(ckpt_path, map_location="cpu")
print("Saved Step:", checkpoint["step"])
print("Validation Loss:", checkpoint["val_loss"])
print("Validation Perplexity:", checkpoint["val_perplexity"])
print("Tokens Processed:", checkpoint["tokens_processed"])

2. Load the Fast Tokenizer

from transformers import PreTrainedTokenizerFast
from huggingface_hub import snapshot_download

tok_dir = snapshot_download(repo_id="anuml/anlp-assignment2-models", allow_patterns="tokenizer/*")
tokenizer = PreTrainedTokenizerFast.from_pretrained(f"{tok_dir}/tokenizer")

tokens = tokenizer.encode("Mixture of Experts pretraining evaluation.")
print("Encoded tokens:", tokens)

πŸ“‚ Repository Directory Layout

.
β”œβ”€β”€ README.md
β”œβ”€β”€ tokenizer/
β”‚   β”œβ”€β”€ vocab.json
β”‚   β”œβ”€β”€ merges.txt
β”‚   β”œβ”€β”€ tokenizer.json
β”‚   └── tokenizer_config.json
β”œβ”€β”€ configs/
β”‚   β”œβ”€β”€ part1_variant[1-5].yaml
β”‚   β”œβ”€β”€ part2_[adamw,cautious,lion,muon,sophia].yaml
β”‚   └── part3_decoding.yaml
β”œβ”€β”€ checkpoints/
β”‚   β”œβ”€β”€ part1/
β”‚   β”‚   β”œβ”€β”€ variant1/model_variant1_best.pt
β”‚   β”‚   β”œβ”€β”€ variant2/model_variant2_best.pt
β”‚   β”‚   β”œβ”€β”€ variant3/model_variant3_best.pt
β”‚   β”‚   β”œβ”€β”€ variant4/model_variant4_best.pt
β”‚   β”‚   └── variant5/model_variant5_best.pt
β”‚   └── part2/
β”‚       β”œβ”€β”€ sweep_adamw_lr/adamw_lr_1e-3_model_best.pt
β”‚       β”œβ”€β”€ sweep_cautious_lr/cautious_lr_1e-3_model_best.pt
β”‚       β”œβ”€β”€ sweep_lion_lr/lion_lr_3e-4_model_best.pt
β”‚       β”œβ”€β”€ sweep_muon_lr/muon_lr_0.01_model_best.pt
β”‚       └── sweep_sophia_lr/sophia_lr_5e-4_model_best.pt
β”œβ”€β”€ results/
β”‚   β”œβ”€β”€ part1/benchmark_results.json
β”‚   β”œβ”€β”€ part1/bleu_scores.json
β”‚   └── part3/decoding_metrics_summary.json
└── assets/
    β”œβ”€β”€ part1/part1_all_expert_heatmaps_grid.png
    β”œβ”€β”€ part2/part2_best_optimizers_comparison_dynamics_comparison.png
    └── part3/part3_decoding_strategies_comparison.png
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Datasets used to train anuml/anlp-assignment2-models