Reinforce Agent playing Pixelcopter-PLE-v0

This is a trained REINFORCE agent playing Pixelcopter-PLE-v0.

Evaluation

  • Mean reward: 24.60
  • Std reward: 14.77
  • Result (mean - std): 9.83

Hyperparameters

  • Hidden size: 64
  • Gamma: 0.99
  • Learning rate: 0.0001
  • Training episodes: 40000
  • Evaluation episodes: 10
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
9.09k params
Tensor type
F32
·
Video Preview
loading

Evaluation results