π0.5 SO-101 vial RLFT smoke checkpoint (step 5)
This is the consolidated RLinf actor checkpoint after the bounded A100 integration run: five PPO updates, four vectorized Isaac environments, and sparse confirmed-placement reward.
It is an RLinf/OpenPI actor full_weights.pt, not a drop-in LeRobot model.safetensors export. Use it with the matching RLinf v0.3 overlay, base model normalization statistics, and training_config.yaml included here.
The run validates rollout, reward, PPO backpropagation, checkpointing, and artifact export. It is not presented as a proven improved policy: evaluation used five stochastic episodes and observed one success.