REINFORCE CartPole-v1

Trained from scratch for Kay Zheng's Hugging Face Deep RL coursework with AI coding assistance. REINFORCE, reward-to-go, batch baseline, 64-unit tanh policy network. Training: 166816 steps with seed 42. Validation seeds 50000–50029. Held-out evaluation: 100 episodes, seeds 100000–100099, greedy policy. Mean reward 497.1, standard deviation 17.15546560137614, mean minus std 479.944534.

Reproduce

Install gymnasium==0.29.1, numpy, torch. Run python train_cartpole.py. policy.pt contains the trained state dictionary. Evaluation rewards are in evaluation.json.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results

  • mean_reward on CartPole-v1
    self-reported
    497.1 +/- 17.15546560137614