Reinforce Agent playing Pixelcopter-PLE-v0

This is a trained REINFORCE policy for PixelCopter from Unit 4 of the Hugging Face Deep RL course.

Best checkpoint: h64_lr1e-4_seed50 at episode 9000.

Replay video uses greedy actions from the learned policy.

Evaluation

  • Greedy mean_reward: 34.90
  • Greedy std_reward: 15.33
  • Greedy score mean-std: 19.57
  • Sample mean_reward: 32.20
  • Sample std_reward: 20.54
  • Sample score mean-std: 11.66
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results