YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

ANLP Assignment 2 checkpoints (lokola13)

Code: the assignment repository on code.iiit.ac.in (src/model, src/part1).

part1/ : Mixture-of-Experts ablation (vi/ja -> en translation)

  • tokenizer.json: the shared 16k byte-level BPE tokenizer used by every run.
  • <run>/model.pt: PyTorch state_dict of src.model.transformer.Transformer, best validation checkpoint.
  • <run>/config.json: the merged YAML config (model section -> TransformerConfig).
  • <run>/train_metrics.json, <run>/test_metrics.json, <run>/training_state.json: training and test results.

Runs: dense, moe_e4_k1, moe_e4_k2, moe_s1_e3_k1, moe_e4_k2_active (each trained for one epoch, 37,965,158 tokens).

Load: model = Transformer(TransformerConfig.from_dict(config["model"])); model.load_state_dict(torch.load("model.pt"))

part2/ : optimizer comparison (pretraining on browndw/human-ai-parallel-corpus)

  • tokenizer.json: the same 16k tokenizer as part1 (the model is part1's dense configuration).
  • <optimizer>/model.pt: PyTorch state_dict of src.model.transformer.Transformer after exactly 1x the training corpus (48,350,976 tokens).
  • <optimizer>/config.json, metrics.json, milestones.json (validation loss and continuation BLEU at every 0.1x of the corpus), training_state.json.

Optimizers (all hand-written torch.optim.Optimizer subclasses, Appendix A of Wen et al. 2025): adamw (Alg. 1), nadamw (Alg. 2), lion (Alg. 3), muon (Alg. 8).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support