PPO LunarLander-v3

This repository contains a Stable-Baselines3 PPO agent trained to solve the Gymnasium LunarLander environment.

Training setup

  • Algorithm: PPO
  • Environment: LunarLander-v3
  • Policy: MlpPolicy
  • Training timesteps: 1,000,000

Evaluation

The agent was evaluated on the LunarLander environment with deterministic rollout settings.

Notes

This model is intended for experimentation and educational purposes.

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support