Reinforce Agent playing Pixelcopter-PLE-v0

This is a trained REINFORCE agent playing Pixelcopter-PLE-v0.

This model was trained as part of the Hugging Face Deep Reinforcement Learning Course Unit 4.

Evaluation

  • Mean reward: 7.20
  • Standard deviation: 8.35
  • Final score (mean - std): -1.15

Hyperparameters

  • Hidden size: 64
  • Training episodes: 5000
  • Evaluation episodes: 10
  • Maximum steps: 10000
  • Gamma: 0.99
  • Learning rate: 0.0001

Course: https://huggingface.co/learn/deep-rl-course/unit4/hands-on

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results