REINFORCE Agent โ€” Pixelcopter-PLE-v0

This model was trained using the REINFORCE Monte Carlo Policy Gradient algorithm.

Environment

Pixelcopter-PLE-v0

Training

Training episodes: 50,000

Hidden size: 64

Learning rate: 0.0001

Gamma: 0.99

Evaluation

Mean reward: 58.20

Standard deviation: 46.93

Certification result:

11.27

The certification result is calculated as:

mean reward - standard deviation

Course

Hugging Face Deep Reinforcement Learning Course

Unit 4 โ€” Policy Gradient Methods

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results