PPO Agent Playing LunarLander-v2
This is a trained model of a Proximal Policy Optimization (PPO) agent playing LunarLander-v2, coded from scratch using PyTorch based on the CleanRL implementation for the Hugging Face Deep Reinforcement Learning Course Unit 8 (Part 1).
Evaluation Results
- Mean Reward: 271.06 +/- 21.00 (evaluated over 10 episodes)
- Environment: LunarLander-v2
Video Replay
Hyperparameters
{
"exp_name": "ppo",
"seed": 1,
"torch_deterministic": True,
"cuda": True,
"track": False,
"wandb_project_name": "cleanRL",
"wandb_entity": None,
"capture_video": False,
"env_id": "LunarLander-v2",
"total_timesteps": 500000,
"learning_rate": 0.00025,
"num_envs": 4,
"num_steps": 128,
"anneal_lr": True,
"gae": True,
"gamma": 0.99,
"gae_lambda": 0.95,
"num_minibatches": 4,
"update_epochs": 4,
"norm_adv": True,
"clip_coef": 0.2,
"clip_vloss": True,
"ent_coef": 0.01,
"vf_coef": 0.5,
"max_grad_norm": 0.5,
"target_kl": None,
"repo_id": "Avinash76812/ppo-LunarLander-v2"
}
Evaluation results
- mean_reward on LunarLander-v2self-reported271.06 +/- 21.00