yx49k 1B multimodal checkpoints

0.98B hybrid gated-delta-attention model with a from-scratch vision tower, trained on a TPU Research Cloud v4-64 (55B tokens: 25B trial + 30B continuation over a calibrated image/text/code/math mix), plus the post-training ladder. Tokenizer: yx49k (49,152, exact subset of Qwen3.5's vocabulary). Full technical report and code: https://github.com/yxanul/yxTPU

directory contents
base-trial-48000 orbax train state, 25B-token trial (post-anneal)
base-cont30b-57220 orbax train state, 55B-token continuation (post-anneal)
sft1-1240 SFT pickle (Mephisto IF+Knowledge)
gold1-4449 GOLD pickle on SFT1 (IFEval mean 40.8, panel 24/32)
sft2-5200 SFT pickle (IF+Knowledge+AMD+MathCode; IFEval mean 55.3)
gold2-4449 GOLD pickle on SFT2 (panel 27/32, gsm8k 4.6)

Pickles are state.pkl TrainStateNNX pure-dicts (weights+optimizer); orbax dirs restore via this repo's CheckpointIO.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support