PPO agent playing LunarLander-v2

This is a trained model of an PPO agent playing LunarLander-v2 for the Hugging Face Deep Reinforcement Learning Course (Unit 8 PI).

Evaluation Results

  • Mean Reward: 250.00 +/- 20.00
  • Environment: LunarLander-v2
  • Algorithm: PPO
  • Library: deep-rl-course

Usage

Trained and evaluated for the Hugging Face Deep RL Course certification.

Downloads last month
17
Video Preview
loading

Evaluation results