PixelCopter REINFORCE
A REINFORCE policy trained using PyTorch on the Pixelcopter-PLE-v0 environment.
Model Details
- Algorithm: REINFORCE
- Environment: Pixelcopter-PLE-v0
- Framework: PyTorch
- Device: CPU
- Observation Space: 7
- Action Space: 2
Training
- Training Episodes: 15,000
- Hidden Size: 64
- Learning Rate: 0.0001
- Gamma: 0.99
- Maximum Steps: 2,000
Evaluation
- Mean Reward: 17.3
- Standard Deviation: 9.40
- Result (Mean - Std): 7.90
Required certification threshold: 5.0
Files
pixelcopter_policy.pth— trained PyTorch policypixelcopter_hyperparameters.json— training hyperparameters
Evaluation results
- mean_reward on Pixelcopter-PLE-v0self-reported17.300
- std_reward on Pixelcopter-PLE-v0self-reported9.400