Reinforce Agent playing CartPole-v1

REINFORCE agent trained as part of the Hugging Face Deep Reinforcement Learning Course Unit 4.

Evaluation

  • Mean reward: 492.30
  • Standard deviation: 20.28
  • Unit 4 score (mean - std): 472.02

Hyperparameters

  • Hidden size: 64
  • Training episodes: 3000
  • Evaluation episodes: 10
  • Maximum steps: 500
  • Gamma: 0.99
  • Learning rate: 0.01
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results