Reinforce Pixelcopter
This model was trained using the REINFORCE policy-gradient algorithm with PyTorch on the Pixelcopter-PLE-v0 environment.
Evaluation
- Mean reward: 28.70
- Standard deviation: 17.02
- Certification score: 11.68
- Evaluation episodes: 10
- Learning rate: 1e-4
- Gamma: 0.99
- Training episodes: approximately 24,000+
The certification score is calculated as:
mean reward - standard deviation = 28.70 - 17.02 = 11.68
Evaluation results
- mean_reward on Pixelcopter-PLE-v0self-reported28.70 +/- 17.02