Teacher exact39K random curriculum Stage3 โ€” GRPO

This repository contains the final, merged GRPO checkpoint of our_stage1_ep2_teacher_exact39k_random42_stage3_exact39k_ep1.

The starting checkpoint is the Teacher exact39K Stage3 model: Stage1 training for 2 epochs, a seed-42 random curriculum of three 13K subsets with 2 epochs per subset, followed by Stage3 training on teacher-correct examples for 1 epoch.

The published checkpoint is GRPO global step 12, the final saved step of the completed run. The four distributed actor shards were merged into four Hugging Face safetensors shards. Tokenizer and image processor files are included.

RL configuration

Setting Value
Algorithm GRPO
Training data Thyme-RL, maximum 3,200 training examples
Epochs 1
GPUs 4
Latent size 10
Rollouts per prompt 8
Sampling temperature 0.5
Learning rate 1e-6
KL coefficient 0.01
Monet RL sigma 10.0
Final saved step 12

The architecture is Qwen2.5-VL. Monet latent reasoning requires the Monet inference/runtime implementation used by the training project.

Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support