REINFORCE Agent playing Pixelcopter-PLE-v0
This is a trained REINFORCE agent for Pixelcopter-PLE-v0, created for Unit 4 of the Hugging Face Deep Reinforcement Learning Course.
Evaluation
- Training episodes: 30000
- Evaluation episodes: 10
- Mean reward: 42.30
- Standard deviation: 20.96
- Course result (
mean_reward - std_reward): 21.34 - Required Unit 4 threshold: 5
Hyperparameters
- Hidden layers: 64 → 128
- Gamma: 0.99
- Learning rate: 1e-4
- Max timesteps per episode: 10000
- Algorithm: REINFORCE / Monte Carlo Policy Gradient
The environment was run using a Gymnasium-compatible adaptation of the original Unit 4 notebook.
Evaluation results
- mean_reward on Pixelcopter-PLE-v0self-reported42.30 +/- 20.96