πŸš€ PPO Agent for LunarLander-v3

This project trains a reinforcement learning agent using Proximal Policy Optimization (PPO) to solve the LunarLander-v3 environment from Gymnasium. The agent learns to land a lunar module safely between two flags using Box2D physics.


🧠 Project Summary

  • Environment: LunarLander-v3 (Box2D)
  • Algorithm: PPO (Stable-Baselines3)
  • Training: 1 million timesteps using vectorized environments
  • Evaluation: Mean reward ~269 Β± 11.6 over 10 episodes
  • Replay: Manually recorded and uploaded to Hugging Face

πŸ“ Hugging Face Model Card

The trained model is available at:

πŸ”— DevilNReality/ppo-LunarLander-v3

Includes:

  • Model weights
  • Evaluation metrics
  • Replay video
  • Auto-generated metadata

🎬 Replay Video

Watch the agent land in the environment:

Click to view replay


βœ… Certification

This model satisfies the Hugging Face Deep RL course Unit 1 requirements:

  • Trained PPO agent on LunarLander-v3
  • Evaluation score β‰₯ 200
  • Model pushed to Hugging Face
  • Replay video included

πŸ“š References

Downloads last month
1
Video Preview
loading

Evaluation results