browndw/human-ai-parallel-corpus
Viewer β’ Updated β’ 66.3k β’ 1.53k β’ 4
This repository hosts the trained model checkpoints, custom tokenizer, configurations, and evaluation outputs for ANLP Assignment 2 (Monsoon 2026).
belumind/en-vi-ja-curated-500k-triplets for 90M tokens (3x baseline budget).browndw/human-ai-parallel-corpus for 30M tokens.All variants share an identical backbone ($D_{model}=512$, $N_{layers}=6$, $N_{heads}=8$, RMSNorm, RoPE, Vocab=16,384) and were trained on an identical 90,000,000 token budget:
| Variant | Description | Total Params | Active Params | Val Loss | Val PPL | Test BLEU | Routing Strategy |
|---|---|---|---|---|---|---|---|
| V1 | Standard Dense 2-Layer MLP | 35.69M | 35.67M | 3.3361 | 28.11 | 0.6585 | Dense Baseline ($d_{ff}=2048$) |
| V2 | MoE (4 experts, Top-1 active) | 54.58M | 35.67M | 3.3980 | 29.90 | 0.6653 | Top-1 softmax routing |
| V3 | MoE (4 experts, Top-2 active) | 54.58M | 41.97M | 3.0185 | 20.46 | 0.7188 | Top-2 normalized routing |
| V4 | Shared + Routed (1 shared + Top-1/3 routed) | 54.58M | 41.97M | 3.0336 | 20.77 | 0.7099 | 1 always-active + Top-1 routed |
| V5 | Parameter-Matched MoE (4 experts, Top-2) | 45.13M | 35.67M | 3.2085 | 24.74 | 0.6865 | Top-2 ($d_{ff}=1024$, matched active FLOPs) |
Comparison of all 5 optimizers on browndw/human-ai-parallel-corpus under an identical 30M token budget:
| Category | Optimizer | Best LR | State Memory | Final Val Loss | Final Val PPL | Final BLEU | Speedup to 3.5 Loss |
|---|---|---|---|---|---|---|---|
| Cat 1: Standard Adaptive | AdamW | $1 \times 10^{-3}$ | $2\times$ (8 B/param) | 3.4309 | 30.90 | 0.6133 | $1.00\times$ (0.9x tokens) |
| Cat 2: Variance-Reduced | Cautious AdamW | $1 \times 10^{-3}$ | $2\times$ (8 B/param) | 3.4209 | 30.60 | 0.7022 | $1.05\times$ (0.8x tokens) |
| Cat 3: Memory-Efficient | Lion | $3 \times 10^{-4}$ | $1\times$ (4 B/param) | 3.6085 | 36.91 | 0.6933 | $0.90\times$ (1.0x tokens) |
| Cat 4: Matrix-Based | Muon | $1 \times 10^{-2}$ | $1\times$ (4 B/param) | 3.3077 | 27.32 | 0.7297 | $\mathbf{1.35\times}$ (0.7x tokens) |
| Cat 5: Hessian-Based (Bonus) | Sophia-G | $5 \times 10^{-4}$ | $2\times$ (8 B/param) | 4.4871 | 88.86 | 0.3965 | $> 1.0\times$ (did not reach) |
huggingface_hub
from huggingface_hub import hf_hub_download
import torch
# Download Variant 3 best checkpoint
ckpt_path = hf_hub_download(
repo_id="anuml/anlp-assignment2-models",
filename="checkpoints/part1/variant3/model_variant3_best.pt"
)
# Load checkpoint state dict & training metadata
checkpoint = torch.load(ckpt_path, map_location="cpu")
print("Saved Step:", checkpoint["step"])
print("Validation Loss:", checkpoint["val_loss"])
print("Validation Perplexity:", checkpoint["val_perplexity"])
print("Tokens Processed:", checkpoint["tokens_processed"])
from transformers import PreTrainedTokenizerFast
from huggingface_hub import snapshot_download
tok_dir = snapshot_download(repo_id="anuml/anlp-assignment2-models", allow_patterns="tokenizer/*")
tokenizer = PreTrainedTokenizerFast.from_pretrained(f"{tok_dir}/tokenizer")
tokens = tokenizer.encode("Mixture of Experts pretraining evaluation.")
print("Encoded tokens:", tokens)
.
βββ README.md
βββ tokenizer/
β βββ vocab.json
β βββ merges.txt
β βββ tokenizer.json
β βββ tokenizer_config.json
βββ configs/
β βββ part1_variant[1-5].yaml
β βββ part2_[adamw,cautious,lion,muon,sophia].yaml
β βββ part3_decoding.yaml
βββ checkpoints/
β βββ part1/
β β βββ variant1/model_variant1_best.pt
β β βββ variant2/model_variant2_best.pt
β β βββ variant3/model_variant3_best.pt
β β βββ variant4/model_variant4_best.pt
β β βββ variant5/model_variant5_best.pt
β βββ part2/
β βββ sweep_adamw_lr/adamw_lr_1e-3_model_best.pt
β βββ sweep_cautious_lr/cautious_lr_1e-3_model_best.pt
β βββ sweep_lion_lr/lion_lr_3e-4_model_best.pt
β βββ sweep_muon_lr/muon_lr_0.01_model_best.pt
β βββ sweep_sophia_lr/sophia_lr_5e-4_model_best.pt
βββ results/
β βββ part1/benchmark_results.json
β βββ part1/bleu_scores.json
β βββ part3/decoding_metrics_summary.json
βββ assets/
βββ part1/part1_all_expert_heatmaps_grid.png
βββ part2/part2_best_optimizers_comparison_dynamics_comparison.png
βββ part3/part3_decoding_strategies_comparison.png