PPO Agent Playing LunarLander-v3

This is a trained model of a PPO agent playing LunarLander-v3.

Hyperparameters

  • exp_name: ppo
  • seed: 1
  • torch_deterministic: True
  • cuda: True
  • env_id: LunarLander-v3
  • total_timesteps: 1000000
  • learning_rate: 0.0003
  • num_envs: 16
  • num_steps: 256
  • anneal_lr: True
  • gae: True
  • gamma: 0.999
  • gae_lambda: 0.98
  • num_minibatches: 16
  • update_epochs: 4
  • norm_adv: True
  • clip_coef: 0.2
  • clip_vloss: True
  • ent_coef: 0.01
  • vf_coef: 0.5
  • max_grad_norm: 0.5
  • target_kl: None
  • repo_id: suveda999/ppo-LunarLander-v3-cleanrl
  • batch_size: 4096
  • minibatch_size: 256
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results