sun1638650145 commited on
Commit
e50ac9b
1 Parent(s): b819b9f

发布强化学习模型到Hugging Face Hub.

Browse files
README.md ADDED
@@ -0,0 +1,62 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ tag:
3
+ - LunarLander-v2
4
+ - ppo
5
+ - deep-reinforcement-learning
6
+ - reinforcement-learning
7
+ - custom-implementation
8
+ - deep-rl-class
9
+ model-index:
10
+ - name: PPO
11
+ results:
12
+ - metrics:
13
+ - type: mean_reward
14
+ value: -121.77 +/- 30.58
15
+ name: mean_reward
16
+ task:
17
+ type: reinforcement-learning
18
+ name: reinforcement-learning
19
+ dataset:
20
+ name: LunarLander-v2
21
+ type: LunarLander-v2
22
+ ---
23
+
24
+ # 使用PPO智能体来玩 LunarLander-v2
25
+
26
+ 这是一个使用PPO训练有素的模型玩 LunarLander-v2.
27
+ 要学习编写你自己的PPO智能体并训练它,
28
+ 请查阅深度强化学习课程第8单元: https://github.com/huggingface/deep-rl-class/tree/main/unit8
29
+
30
+ # 超参数
31
+ ```python
32
+ {'exp_name': 'ppo'
33
+ 'seed': 1
34
+ 'torch_deterministic': True
35
+ 'cuda': True
36
+ 'track': False
37
+ 'wandb_project_name': 'cleanRL'
38
+ 'wandb_entity': None
39
+ 'capture_video': False
40
+ 'env_id': 'LunarLander-v2'
41
+ 'total_timesteps': 50000
42
+ 'learning_rate': 0.00025
43
+ 'num_envs': 4
44
+ 'num_steps': 128
45
+ 'anneal_lr': True
46
+ 'gae': True
47
+ 'gamma': 0.99
48
+ 'gae_lambda': 0.95
49
+ 'num_minibatches': 4
50
+ 'update_epochs': 4
51
+ 'norm_adv': True
52
+ 'clip_coef': 0.2
53
+ 'clip_vloss': True
54
+ 'ent_coef': 0.01
55
+ 'vf_coef': 0.5
56
+ 'max_grad_norm': 0.5
57
+ 'target_kl': None
58
+ 'repo_id': 'sun1638650145/PyTorch-PPO-LunarLander-v2'
59
+ 'batch_size': 512
60
+ 'minibatch_size': 128}
61
+ ```
62
+
hyperparameters.json ADDED
@@ -0,0 +1 @@
 
 
1
+ {"env_id": "LunarLander-v2", "mean_reward": -121.77415005450985, "std_reward": 30.583650538281326, "n_evaluation_episodes": 10, "eval_datetime": "2022-08-12T10:48:22.717841"}
logs/events.out.tfevents.1660272473.sunruiqideMacBook-Pro-2020.local.37073.0 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bce4e19a80851316e74b2fc987db898f1891021ec60ce3bb75cb6fc3ee9d5687
3
+ size 111642
model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6127d2d7475300d0900eda3d3b930b772c6c758e4f3b663c7ca2c17e27a9667c
3
+ size 42689
replay.mp4 ADDED
Binary file (53 kB). View file