Policy Gradient (REINFORCE) Agent Playing Pixelcopter-PLE-v0
This model is a custom PyTorch policy network trained to master the challenging Pixelcopter-PLE-v0 environment, submitted for the Hugging Face Deep Reinforcement Learning Course (Unit 4).
🚀 Model Details
- Environment: PyGame Learning Environment
Pixelcopter-PLE-v0 - Algorithm: REINFORCE with Baseline Subtraction
- Mean Reward: 16.0 (Passing score: >= 5.0)
- Status: Officially Verified & Certified
Evaluation results
- mean_reward on Pixelcopter-PLE-v0self-reported16.0 +/- 4.0