REINFORCE (Monte Carlo Policy Gradient) Agent playing CartPole-v1

This is a trained model of a REINFORCE policy network playing CartPole-v1. Built and trained in PyTorch with variance reduction baseline normalization.

Model Architecture

  • Input: 4 continuous state features (Cart Position, Cart Velocity, Pole Angle, Pole Angular Velocity)
  • Hidden Layers: Linear(4, 64) -> ReLU -> Linear(64, 64) -> ReLU -> Linear(64, 2)
  • Action Selection: Categorical distribution over discrete actions [Push cart left, Push cart right]

Evaluation Results

  • Mean Reward: 9.33 +/- 0.81 (Passing threshold: $\ge 350.0$)
  • Number of Evaluation Episodes: 100
  • Model weights: Saved in reinforce_cartpole.pt

Trained and submitted by Subhash3008.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results