talkie-web-13b-it-dpo (partial checkpoint)
Stage-2 online DPO checkpoint of the Talkie 13B web model, preference-trained
on top of the Stage-1 SFT checkpoint with a Claude (claude-sonnet-4-6) pairwise judge.
Note: This is a partial / intermediate checkpoint at step 1500 of 4000 (uploaded mid-run). It is not the final model.
- Base:
talkie-web-13b-base(IT vocab, 65540 tokens) - Objective: DPO (beta=0.05, EMA reference ref_ema=0.1), lr 1e-5 cosine
- Architecture: 40-layer GPT (RoPE, RMSNorm, SwiGLU, HeadGain/WeightGain/ActGain)
Files
step_1500.ptโ{"model_state_dict": ...}, bf16vocab.txtโ tiktoken BPE vocab (IT)
Load with the talkie library's load_checkpoint (target_vocab_size=65540).
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support
Model tree for Jiminator/talkie-web-13b-it-dpo
Base model
talkie-lm/talkie-web-13b-base