PPO LunarLander-v2

Evaluation

Mean reward: 263.21 +/- 17.22

Score used by the Deep RL Course certification checker:

263.21 - 17.22 = 245.99

Important note

The PPO policy was originally trained using LunarLander-v3.

It was subsequently evaluated in the legacy LunarLander-v2 environment using Gymnasium 0.29.1 for compatibility with the legacy Hugging Face Deep RL Course certification checker.

The model therefore uses the same trained policy while the reported certification evaluation is a genuine evaluation on LunarLander-v2.

Model

  • Algorithm: PPO
  • Library: Stable-Baselines3
  • Environment evaluation: LunarLander-v2
  • Evaluation episodes: 10
Downloads last month
10
Video Preview
loading

Evaluation results