PPO Agent Playing LunarLander-v2
This is a trained model of a PPO agent playing LunarLander-v2, implemented from scratch following the CleanRL tutorial.
Result: 266.10 +/- 15.82
Hyperparameters
{
'gym_id': 'LunarLander-v2',
'seed': 1,
'num_envs': 16,
'num_steps': 1024,
'total_timesteps': 10000000,
'capture_video': False,
'torch_deterministic': True,
'cuda': True,
'learning_rate': 0.00025,
'anneal_lr': True,
'gae': True,
'gamma': 0.999,
'gae_lambda': 0.98,
'num_minibatches': 16,
'update_epochs': 4,
'norm_adv': True,
'clip_coef': 0.2,
'clip_vloss': True,
'ent_coef': 0.003,
'vf_coef': 0.5,
'max_grad_norm': 0.5,
'target_kl': None,
'batch_size': 16384,
'minibatch_size': 1024,
'env_id': 'LunarLander-v2',
}
Evaluation results
- mean_reward on LunarLander-v2self-reported266.10 +/- 15.82