PPO Agent Playing LunarLander-v2

This is a trained model of a Proximal Policy Optimization (PPO) agent playing LunarLander-v2, coded from scratch using PyTorch based on the CleanRL implementation for the Hugging Face Deep Reinforcement Learning Course Unit 8 (Part 1).

Evaluation Results

  • Mean Reward: 271.06 +/- 21.00 (evaluated over 10 episodes)
  • Environment: LunarLander-v2

Video Replay

Replay

Hyperparameters

{
    "exp_name": "ppo",
    "seed": 1,
    "torch_deterministic": True,
    "cuda": True,
    "track": False,
    "wandb_project_name": "cleanRL",
    "wandb_entity": None,
    "capture_video": False,
    "env_id": "LunarLander-v2",
    "total_timesteps": 500000,
    "learning_rate": 0.00025,
    "num_envs": 4,
    "num_steps": 128,
    "anneal_lr": True,
    "gae": True,
    "gamma": 0.99,
    "gae_lambda": 0.95,
    "num_minibatches": 4,
    "update_epochs": 4,
    "norm_adv": True,
    "clip_coef": 0.2,
    "clip_vloss": True,
    "ent_coef": 0.01,
    "vf_coef": 0.5,
    "max_grad_norm": 0.5,
    "target_kl": None,
    "repo_id": "Avinash76812/ppo-LunarLander-v2"
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results