scale_base_blend_dr @ step 400

RL checkpoint. Base TikZilla-3B (SFT), trained with Dr.GRPO/DAPO against the DSV blend_dr verifiable reward (0.7路DSV + 0.3路RenderGraph, both symbolic verifiers -- no learned reward model).

  • data: explcre/datikz_v4_rl_hybrid_428k (427,553 rows, 46.6% Kimi-K3 descriptions / 53.4% DaTikZ-v4 vlm_description, tagged per row)
  • verifier: explcre/DeTikZify branch dsv-verifier @ c348406
  • config: batch 84, rollout n=8, max_response_length 2048, 1500 steps planned

Evaluate with the matching input distribution. The same checkpoint scores differently on vlmdesc984 vs kimidesc984, and differently again under different TeX/compile-gate versions -- always report the measurement config with the number.

Downloads last month
16
Safetensors
Model size
3B params
Tensor type
BF16
Video Preview
loading