YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

PPO LunarLander-v3

使用 PPO 算法训练 LunarLander-v3 智能体。

实验环境

  • Python 3.11.9
  • Gymnasium 1.4.0
  • Stable-Baselines3 2.9.0
  • PyTorch 2.4.1
  • 训练设备:CPU

训练参数

  • 算法:PPO
  • 策略:MlpPolicy
  • 并行环境数量:16
  • 训练步数:1,000,000

实验结果

  • 平均奖励:254.57
  • 标准差:18.96
  • 结论:智能体已经学会基本完成着陆任务。

文件说明

  • ppo-lunarlander-v3.zip:训练好的 PPO 模型
  • videos/:模型运行视频
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support