Frozen-expert MoE router

Small MLP gate over five frozen 99M squares64 specialists. Only this router is trained; the experts stay in their own repos.

This file is latest.pt at router step 550 (2026-09-14 19:08 UTC). Train loss ~0.6869. Val source-cls CE ~0.4296. Val acc ~0.839.

Not a 99M policy. Not the incumbent, puzzle, endgame, opening, or middlegame expert.

Experts

Architecture

  • Stem: frozen incumbent global_hidden (736d)
  • Head: LayerNorm β†’ Linear(736,256) β†’ GELU β†’ Linear(256,256) β†’ GELU β†’ Linear(256,5)
  • Params: 257,221
  • Vocab: compact 1968
  • Dispatch: hard argmax, one specialist forward (incumbent encode is reused if chosen)

Files

  • latest.pt β€” router weights + expert repo list
  • step_000550.pt
  • router_config.json
  • train.log
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support