CleanRL PPO Agent Playing LunarLander-v2

This is a trained model of a PPO agent playing LunarLander-v2 using the CleanRL implementation with PyTorch.

Evaluation Results

  • Mean Reward: -182.76 +/- 119.89 (evaluated over 10 episodes)

Hyperparameters

  • exp_name: cleanrl_ppo
  • seed: 1
  • torch_deterministic: True
  • cuda: True
  • track: False
  • wandb_project_name: cleanRL
  • wandb_entity: None
  • capture_video: False
  • env_id: LunarLander-v2
  • total_timesteps: 50000
  • learning_rate: 0.00025
  • num_envs: 8
  • num_steps: 128
  • anneal_lr: True
  • gae: True
  • gamma: 0.99
  • gae_lambda: 0.95
  • num_minibatches: 4
  • update_epochs: 4
  • norm_adv: True
  • clip_coef: 0.2
  • clip_vloss: True
  • ent_coef: 0.01
  • vf_coef: 0.5
  • max_grad_norm: 0.5
  • target_kl: None
  • repo_id: Atharva1232/cleanrl-ppo-LunarLander-v2
  • batch_size: 1024
  • minibatch_size: 256

To learn more check Unit 8 of the Deep Reinforcement Learning Course: https://huggingface.co/deep-rl-course/unit8/introduction

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results