Reinforce Agent playing CartPole-v1
This is a trained Reinforce agent for Unit 4 of the Hugging Face Deep Reinforcement Learning Course.
Evaluation
- Environment: CartPole-v1
- Mean reward: 500.00
- Standard deviation: 0.00
- Course result: 500.00
- Required result: 350
- Evaluation episodes: 10
Training
- Algorithm: REINFORCE / Monte Carlo Policy Gradient
- Hidden layer: 16
- Gamma: 1.0
- Learning rate: 1e-2
- Training episodes: 1500
Evaluation results
- mean_reward on CartPole-v1self-reported500.00 +/- 0.00