Reinforcement Learning
stable-baselines3
English
LunarLander-v3
deep-reinforcement-learning
Eval Results (legacy)
Instructions to use DevilNReality/ppo-LunarLander-v3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- stable-baselines3
How to use DevilNReality/ppo-LunarLander-v3 with stable-baselines3:
from huggingface_sb3 import load_from_hub checkpoint = load_from_hub( repo_id="DevilNReality/ppo-LunarLander-v3", filename="{MODEL FILENAME}.zip", ) - Notebooks
- Google Colab
- Kaggle
π PPO Agent for LunarLander-v3
This project trains a reinforcement learning agent using Proximal Policy Optimization (PPO) to solve the LunarLander-v3 environment from Gymnasium. The agent learns to land a lunar module safely between two flags using Box2D physics.
π§ Project Summary
- Environment: LunarLander-v3 (Box2D)
- Algorithm: PPO (Stable-Baselines3)
- Training: 1 million timesteps using vectorized environments
- Evaluation: Mean reward ~269 Β± 11.6 over 10 episodes
- Replay: Manually recorded and uploaded to Hugging Face
π Hugging Face Model Card
The trained model is available at:
π DevilNReality/ppo-LunarLander-v3
Includes:
- Model weights
- Evaluation metrics
- Replay video
- Auto-generated metadata
π¬ Replay Video
Watch the agent land in the environment:
β Certification
This model satisfies the Hugging Face Deep RL course Unit 1 requirements:
- Trained PPO agent on
LunarLander-v3 - Evaluation score β₯ 200
- Model pushed to Hugging Face
- Replay video included
π References
- Downloads last month
- 1
Evaluation results
- mean_reward on LunarLander-v3self-reported269.15 +/- 11.63