PPO LunarLander-v2

Kay Zheng's Unit 1 coursework, trained from scratch with AI coding assistance. Stable Baselines3 2.3.2, Gymnasium 0.29.1, box2d-py 2.3.8, seed 42. Training: 761856 steps on 8 environments, local CPU. Validation seeds 50000–50019; final independent evaluation seeds 100000–100099. Mean reward 248.60384153527525, standard deviation 22.2469236328401; mean minus std 226.356918. All evaluation episode rewards are included in evaluation.json.

Reproduce with python train_lunar.py. Load with PPO.load('model.zip').

Downloads last month
8
Video Preview
loading

Evaluation results

  • mean_reward on LunarLander-v2
    self-reported
    248.60384153527525 +/- 22.2469236328401