PixelCopter REINFORCE

A REINFORCE policy trained using PyTorch on the Pixelcopter-PLE-v0 environment.

Model Details

  • Algorithm: REINFORCE
  • Environment: Pixelcopter-PLE-v0
  • Framework: PyTorch
  • Device: CPU
  • Observation Space: 7
  • Action Space: 2

Training

  • Training Episodes: 15,000
  • Hidden Size: 64
  • Learning Rate: 0.0001
  • Gamma: 0.99
  • Maximum Steps: 2,000

Evaluation

  • Mean Reward: 17.3
  • Standard Deviation: 9.40
  • Result (Mean - Std): 7.90

Required certification threshold: 5.0

Files

  • pixelcopter_policy.pth — trained PyTorch policy
  • pixelcopter_hyperparameters.json — training hyperparameters
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results