a2-task1-v3_moe_top2

ANLP Assignment 2, Task 1 (Mixture of Experts): Vietnamese/Japanese to English translation, trained from scratch on belumind/en-vi-ja-curated-500k-triplets.

FFN variant: MoE, 4 experts (width 512), top-2 routing. Total params matched to the dense model.

Decoder-only transformer: d_model 512, 8 layers, 8 heads, context 256, learned absolute positions, pre-LayerNorm, bias-free GELU FFN, 32k BPE vocabulary with the embedding tied to the LM head.

total params active params train tokens test ppl ppl vi->en ppl ja->en BLEU BLEU vi->en BLEU ja->en chrF
41,714,688 33,326,080 73.0M 4.946 4.188 5.841 37.68 42.80 32.46 59.09

Input layout: [bos] <2en> source [sep] target [eos] (loss on the target only); greedy decoding.

Files: model.pt (a torch.save dict: model state_dict, step, tokens_seen, config), tokenizer.json (HF tokenizers), config.yaml (training config), metrics.json (full eval output).

Loading (with the assignment code on the path):

import torch
from src.part1.config import TransformerConfig, resolve_ffn_dims
from src.part1.model import build_model

ckpt = torch.load('model.pt', map_location='cpu', weights_only=False)
model = build_model(resolve_ffn_dims(TransformerConfig(**ckpt['config'])))
model.load_state_dict(ckpt['model'])
Downloads last month
6
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train FrenchKnuckles/a2-task1-v3_moe_top2