PPO agent playing ML-Agents-Pyramids

This is a trained model of an PPO agent playing ML-Agents-Pyramids for the Hugging Face Deep Reinforcement Learning Course (Unit 5 P2).

Evaluation Results

  • Mean Reward: 1.85 +/- 0.30
  • Environment: ML-Agents-Pyramids
  • Algorithm: PPO
  • Library: ml-agents

Usage

Trained and evaluated for the Hugging Face Deep RL Course certification.

Downloads last month
12
Video Preview
loading

Evaluation results