SAC PandaReachDense-v3

Kay Zheng's Unit 6 coursework, trained from scratch with AI coding assistance. Stable Baselines3 2.3.2, Gymnasium 0.29.1, panda-gym 3.0.7, pybullet 3.2.7, seed 42. Training: 20000 environment steps, CPU. Held-out evaluation: 100 episodes, seeds 100000–100099. Mean reward -0.2263398146163672, std 0.1048106702883708, mean minus std -0.331150. Full evaluation rewards are in evaluation.json.

Reproduce with python train_panda.py; load with SAC.load('model.zip'). On current macOS SDK, building pybullet 3.2.7 required CFLAGS='-Dfdopen=fdopen' to prevent an obsolete bundled zlib macro from shadowing the system declaration.

Downloads last month
14
Video Preview
loading

Evaluation results

  • mean_reward on PandaReachDense-v3
    self-reported
    -0.2263398146163672 +/- 0.1048106702883708