talkie-web-13b-it-dpo (partial checkpoint)

Stage-2 online DPO checkpoint of the Talkie 13B web model, preference-trained on top of the Stage-1 SFT checkpoint with a Claude (claude-sonnet-4-6) pairwise judge.

Note: This is a partial / intermediate checkpoint at step 1500 of 4000 (uploaded mid-run). It is not the final model.

  • Base: talkie-web-13b-base (IT vocab, 65540 tokens)
  • Objective: DPO (beta=0.05, EMA reference ref_ema=0.1), lr 1e-5 cosine
  • Architecture: 40-layer GPT (RoPE, RMSNorm, SwiGLU, HeadGain/WeightGain/ActGain)

Files

  • step_1500.pt โ€” {"model_state_dict": ...}, bf16
  • vocab.txt โ€” tiktoken BPE vocab (IT)

Load with the talkie library's load_checkpoint (target_vocab_size=65540).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Jiminator/talkie-web-13b-it-dpo

Finetuned
(5)
this model