mohjkhan/anlp-a2

Checkpoints for ANLP Assignment 2 (Mixture of Experts, Optimizers, Decoding Strategies). Each checkpoint is a plain torch.save dict with keys model (state_dict) and config (the TransformerConfig). Load with the training repo's Transformer class; no transformers/AutoModel wrapper is provided.

Part 1 β€” FFN / MoE variants

Decoder-only transformer (dim 512, 8 layers, 8 heads, SwiGLU, 32k vocab, ctx 512, RoPE) trained on belumind/en-vi-ja-curated-500k-triplets (VI+JA→EN) at a matched 30M-token budget.

Part 2 β€” Optimizers

Same dense model pretrained on browndw/human-ai-parallel-corpus to ~1Γ— Chinchilla (β‰ˆ38.5M tokens). Final values:

Optimizer Val loss Val ppl Test BLEU
adamw 4.53 92.7 0.703
mars 4.62 101.1 0.768
adafactor 5.80 331.0 0.487
muon 4.62 101.9 0.774
sophia 4.94 140.0 0.532

Part 3 β€” Decoding

Decoding strategies (greedy, top-k, top-p, beam 1/2/4) were evaluated on EleutherAI/pythia-160m with hamishivi/ROCStories and do not produce trained checkpoints (see the code repo + W&B for results).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support