PPO Agent Playing LunarLander-v2

This is a trained model of a PPO agent playing LunarLander-v2 using PyTorch CleanRL implementation.

Evaluation Results

Mean Reward: 280.50 +/- 15.20 Result Score (Mean - Std): 265.30

Hyperparameters

learning_rate: 0.00025
num_envs: 4
num_steps: 128
total_timesteps: 300000
gamma: 0.99
gae_lambda: 0.95
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results