REINFORCE Agent playing CartPole-v1

This model implements the REINFORCE policy-gradient algorithm using PyTorch.

Environment

CartPole-v1

Evaluation

  • Mean reward: 424.41
  • Standard deviation: 47.23
  • Certification score: 377.18
  • Evaluation episodes: 100
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results