a2-task1-v3_moe_top2
ANLP Assignment 2, Task 1 (Mixture of Experts): Vietnamese/Japanese to English translation, trained from scratch on belumind/en-vi-ja-curated-500k-triplets.
FFN variant: MoE, 4 experts (width 512), top-2 routing. Total params matched to the dense model.
Decoder-only transformer: d_model 512, 8 layers, 8 heads, context 256, learned absolute positions, pre-LayerNorm, bias-free GELU FFN, 32k BPE vocabulary with the embedding tied to the LM head.
| total params | active params | train tokens | test ppl | ppl vi->en | ppl ja->en | BLEU | BLEU vi->en | BLEU ja->en | chrF |
|---|---|---|---|---|---|---|---|---|---|
| 41,714,688 | 33,326,080 | 73.0M | 4.946 | 4.188 | 5.841 | 37.68 | 42.80 | 32.46 | 59.09 |
Input layout: [bos] <2en> source [sep] target [eos] (loss on the target only); greedy decoding.
Files: model.pt (a torch.save dict: model state_dict, step, tokens_seen, config), tokenizer.json (HF tokenizers), config.yaml (training config), metrics.json (full eval output).
Loading (with the assignment code on the path):
import torch
from src.part1.config import TransformerConfig, resolve_ffn_dims
from src.part1.model import build_model
ckpt = torch.load('model.pt', map_location='cpu', weights_only=False)
model = build_model(resolve_ffn_dims(TransformerConfig(**ckpt['config'])))
model.load_state_dict(ckpt['model'])
- Downloads last month
- 6