REINFORCE Agent playing Pixelcopter-PLE-v0

This is a trained REINFORCE agent for Pixelcopter-PLE-v0, created for Unit 4 of the Hugging Face Deep Reinforcement Learning Course.

Evaluation

  • Training episodes: 30000
  • Evaluation episodes: 10
  • Mean reward: 42.30
  • Standard deviation: 20.96
  • Course result (mean_reward - std_reward): 21.34
  • Required Unit 4 threshold: 5

Hyperparameters

  • Hidden layers: 64 → 128
  • Gamma: 0.99
  • Learning rate: 1e-4
  • Max timesteps per episode: 10000
  • Algorithm: REINFORCE / Monte Carlo Policy Gradient

The environment was run using a Gymnasium-compatible adaptation of the original Unit 4 notebook.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results