ppo Agent playing Pyramids

This is a trained model of a ppo agent playing Pyramids, trained with Unity ML-Agents as part of the Hugging Face Deep Reinforcement Learning Course.

Evaluation

Real inference run with the exported ONNX policy in the Unity environment.

  • episodes: 60
  • mean_reward: -1.000
  • std_reward: 0.000
  • score (mean - std): -1.000
Downloads last month
-
Video Preview
loading

Evaluation results