scale_base_blend_dr @ step 400
RL checkpoint. Base TikZilla-3B (SFT), trained with Dr.GRPO/DAPO
against the DSV blend_dr verifiable reward (0.7路DSV + 0.3路RenderGraph, both symbolic verifiers --
no learned reward model).
- data:
explcre/datikz_v4_rl_hybrid_428k(427,553 rows, 46.6% Kimi-K3 descriptions / 53.4% DaTikZ-v4vlm_description, tagged per row) - verifier:
explcre/DeTikZifybranchdsv-verifier@c348406 - config: batch 84, rollout n=8, max_response_length 2048, 1500 steps planned
Evaluate with the matching input distribution. The same checkpoint scores differently on
vlmdesc984 vs kimidesc984, and differently again under different TeX/compile-gate versions --
always report the measurement config with the number.
- Downloads last month
- 16