Reinforce Pixelcopter

This model was trained using the REINFORCE policy-gradient algorithm with PyTorch on the Pixelcopter-PLE-v0 environment.

Evaluation

  • Mean reward: 28.70
  • Standard deviation: 17.02
  • Certification score: 11.68
  • Evaluation episodes: 10
  • Learning rate: 1e-4
  • Gamma: 0.99
  • Training episodes: approximately 24,000+

The certification score is calculated as:

mean reward - standard deviation = 28.70 - 17.02 = 11.68

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results