Reinforce Agent playing Pixelcopter-PLE-v0
This is a trained REINFORCE agent playing Pixelcopter-PLE-v0.
Environment
- Environment: Pixelcopter-PLE-v0
- State space: 7
- Action space: 2
Training
- Training episodes: 20000
- Hidden size: 64
- Gamma: 0.99
- Learning rate: 0.0001
Evaluation
- Mean reward: 20.90
- Standard deviation: 13.33
- Mean - Std: 7.57
This model was created as part of the Hugging Face Deep Reinforcement Learning Course, Unit 4.
Evaluation results
- mean_reward on Pixelcopter-PLE-v0self-reported20.90 +/- 13.33