A2C Agent playing PandaReachDense-v3

This is a trained model of an A2C agent playing PandaReachDense-v3 using the stable-baselines3 library.

  • Mean Reward: -0.24 +/- 0.14
  • Leaderboard Score (Mean - Std): -0.38
  • Certification Requirement: >= -3.5 (Passed!)
Downloads last month
7
Video Preview
loading

Evaluation results