๐ Double DQN ยท LunarLander-v3
LunarLander-v3 ๋ฅผ 1000 ์ํผ์๋ ํ์ต์ํจ Double DQN ์์ด์ ํธ์
๋๋ค.
stable-baselines3 ๊ฐ์ RL ํ๋ ์์ํฌ๋ ์ฐ์ง ์๊ณ , Q๋คํธ์ํฌ / ๋ฆฌํ๋ ์ด ๋ฒํผ /
ํ๊น ๋คํธ์ํฌ๋ฅผ PyTorch ์ฝ๋๋ก ์ง์ ๋ค๋ฃน๋๋ค.
A Double DQN agent for LunarLander-v3, trained for 1000 episodes.
No RL framework (stable-baselines3 etc.) โ the Q-network, replay buffer and target
network are plain PyTorch code.
์ถ์ฒ / Provenance
์ ํํ ๋ฐํ๋๋ค. ์ด ์ฝ๋๋ ๋ฐฑ์ง์์ ์๋ก ์ด ๊ฒ์ด ์๋๋๋ค.
- ๊ธฐ๋ฐ: ๋๋ฆฌ ์ฐ์ด๋ DQN ํํ ๋ฆฌ์ผ ๊ตฌํ(Udacity ์คํ์ผ)์์ ์ถ๋ฐํ Colab ๋ ธํธ๋ถ ์ฝ๋
- ์ฌ๊ธฐ์ ๋ํ ๊ฒ: ์คํ์ด ์ ๋๋ import ๋๋ฝ ์์ , Double DQN ์ ํ, Huber ์์ค + ๊ทธ๋๋์ธํธ ํด๋ฆฌํ, ์๋์ธต 64โ128, ํ๊น๋ง ์ด๊ธฐํ ์ ๋ฆฌ, ์ฒดํฌํฌ์ธํธ ํฌ๋งท, ํ๊ฐ/๋ นํ ์คํฌ๋ฆฝํธ, ์ค์๊ฐ ์น ๋์๋ณด๋, ์ฐฉ๋ฅ ํ์ง ์ ์
- ์์ฑ: ์ ์์ ยทํ์ฅ๊ณผ ๋ฌธ์ํ๋ Claude (Claude Code) ์ ๋์์ ๋ฐ์ ์งํํ์ต๋๋ค.
Honest attribution: this is not a from-scratch implementation. It starts from a widely circulated DQN tutorial notebook (Udacity-style), with bug fixes and the additions listed above; that work and this model card were done with the help of Claude (Claude Code).
์ฑ๋ฅ / Results
์๋ ์์น๋ ํ์ต์ด ๋๋ ๋ชจ๋ธ์ 100๋ฒ ์ฐฉ๋ฅ์์ผ ์ฐ ํ๊ฐ ๊ฒฐ๊ณผ์
๋๋ค
(ํ์ต ์์ฒด๋ 1000 ์ํผ์๋ ยท ํํ ์์ ฮต=0 ยท ๊ณ ์ ์๋, evaluate.py --episodes 100).
Evaluation over 100 episodes of the finished agent (training itself ran for 1000 episodes).
| ์ฒดํฌํฌ์ธํธ | ํ๊ท ๋ณด์ | ์ฐฉ๋ฅ ์ฑ๊ณต๋ฅ | ํ๊ท ์ฐฉ๋ฅ ํ์ง |
|---|---|---|---|
lunarlander_dqn.pth (๊ธฐ๋ณธ) |
223.14 ยฑ 84.58 | 74% | 61.3 / 100 |
lunarlander_dqn_softlanding.pth |
203.30 ยฑ 62.41 | 84% | 64.5 / 100 |
- ํ๊ท ๋ณด์์ด ๋์ ์ชฝ์ ๊ธฐ๋ณธ ๋ชจ๋ธ, ๊น๋ํ๊ฒ ์ฐฉ์งํ๋ ๋น๋๋ soft-landing ์ชฝ์ด ๋์ต๋๋ค.
- ํ์ต 1000ํ ์ค 498ํ์งธ์ 200์ ๊ธฐ์ค์ ๋ํํ๊ณ , ์ดํ ๊ณ์ ํ์ตํด 1000ํ์ ์ฑ์ ์ต๋๋ค.
- ํ์ต ์ค(ฮต=0.05) ์ต๊ทผ 100ํ ์ด๋ํ๊ท ์ต๊ณ ๊ธฐ๋ก: 260.5์ , ๊ทธ ๊ตฌ๊ฐ ์ฐฉ๋ฅ ์ฑ๊ณต๋ฅ 97~98% (์ด๋ํ๊ท ์ "100"์ ํ๊ท ์ ๋ด๋ ์ฐฝ ํฌ๊ธฐ์ด์ง, ํ์ต ํ์๊ฐ ์๋๋๋ค.)
The default checkpoint scores higher on mean reward; the
softlandingone lands cleanly more often (its failures are mostly hovering until timeout rather than crashes).
์ฌ์ฉ๋ฒ / Usage
import gymnasium as gym, torch
from huggingface_hub import hf_hub_download
from dqn_agent import Agent # ์ด ์ ์ฅ์์ dqn_agent.py
ckpt = hf_hub_download("SionJang/lunarlander-v3-double-dqn", "lunarlander_dqn.pth")
env = gym.make("LunarLander-v3", render_mode="human")
agent = Agent(env.observation_space.shape[0], int(env.action_space.n)).load(ckpt)
state, _ = env.reset()
done = False
while not done:
action = agent.act(state, 0.0) # eps=0 โ ํํ ์์ด ์ค๋ ฅ๋ง
state, reward, terminated, truncated, _ = env.step(action)
done = terminated or truncated
env.close()
์ ์ฅ์์ ๋ค์ด ์๋ ์คํฌ๋ฆฝํธ๋ก ๋ฐ๋ก ๋๋ฆด ์๋ ์์ต๋๋ค:
python evaluate.py --ckpt lunarlander_dqn.pth --episodes 100 # ํ๊ฐ
python watch.py --ckpt lunarlander_dqn.pth # ์น์ผ๋ก ๊ฐ์
python record_gif.py --ckpt lunarlander_dqn.pth # GIF ๋
นํ
python train.py --episodes 1000 # ์ฒ์๋ถํฐ ์ฌํ์ต
ํ์ต ์ค์ / Training setup
| ํญ๋ชฉ | ๊ฐ |
|---|---|
| ์๊ณ ๋ฆฌ์ฆ | Double DQN (ํ๋ ์ ํ = ํ์ต๋ง, ๊ฐ์น ํ๊ฐ = ํ๊น๋ง) |
| ์ ๊ฒฝ๋ง | 8 โ 128 โ 128 โ 4 (ReLU) |
| ์ตํฐ๋ง์ด์ | Adam, lr 5e-4 |
| ์์ค | Huber (smooth L1), grad clip 10 |
| ํํ ฮต | 1.0 โ 0.05, ํ๋ง๋ค ร0.995 |
| ์ํผ์๋ | 1000 (ํ๋น ์ต๋ 1000 ์คํ ) |
| ๋ฆฌํ๋ ์ด ๋ฒํผ | 100,000 / ๋ฐฐ์น 64 / 4์คํ ๋ง๋ค ํ์ต |
| ํ๊น๋ง | soft update ฯ = 0.001 |
| ๊ฐ๋ง | 0.99 |
| ํ์ต ์๊ฐ | ์ฝ 28๋ถ (Intel i7-11800H, CPU๋ง ์ฌ์ฉ / CUDA ๋ฏธ์ฌ์ฉ) |
์ฐฉ๋ฅ ํ์ง ์ ์ (Landing style score)
ํ๊ฒฝ ๋ณด์์ "์ฑ๊ณตํ๋๊ฐ"๋ง ๋ณด๊ธฐ ๋๋ฌธ์, "์ผ๋ง๋ ๊น๋ํ๊ฒ ๋ด๋ ธ๋๊ฐ"๋ฅผ ๋ฐ๋ก 0~100์ผ๋ก ๋งค๊ฒผ์ต๋๋ค
(style_score.py). The env reward only tells you whether it landed; this extra score measures how well.
| ํญ๋ชฉ | ๋ฐฐ์ | ๊ธฐ์ค |
|---|---|---|
| ์ค์ ์ ํ๋ | 40 | ์ฐฉ๋ฅ๋ ์ ์ค์์ ๊ฐ๊น์ธ์๋ก |
| ์ ์ง ๋ถ๋๋ฌ์ | 25 | ์ ์ง ์๋๊ฐ ๋๋ฆด์๋ก |
| ์์ธ ์ํ | 20 | ๊ธฐ์ธ๊ธฐยทํ์ ์ด ์ ์์๋ก |
| ์ฐ๋ฃ ์ ์ฝ | 15 | ์์ง ์ ํ ๋น์จ์ด ๋ฎ์์๋ก |
๋ฑ๊ธ / grades: S(90+) A(80+) B(65+) C(45+) D
์ค์๊ฐ ํ์ต ๋์๋ณด๋ / Live training dashboard
train.py๋ ํ์ต ๊ณผ์ ์ ๋ธ๋ผ์ฐ์ ์์ ์ค์๊ฐ์ผ๋ก ๋ณด์ฌ์ฃผ๋ ์น ๋์๋ณด๋๋ฅผ ํจ๊ป ๋์๋๋ค
(http://127.0.0.1:7000). Flask ๊ฐ์ ์น ํ๋ ์์ํฌ ์์ด ํ์ด์ฌ ํ์ค http.server ๋ก
๋์๊ฐ๊ณ , ํ๋ฉด ํ๋ ์์ JPEG๋ก ๋ฐ๊พธ๋ ๋ฐ๋ง Pillow๋ฅผ ์๋๋ค. ์ค์๊ฐ ์ฐฉ๋ฅ ํ๋ฉด + HUD,
ํ์ต ๊ณก์ , ฮต ๊ฐ์ ๊ณก์ , ์ฐฉ๋ฅ ํ์ง ์ ์, ์๋ ์ฐฉ๋ฅ GIF ๋ชจ์์ด ๋ค์ด ์์ต๋๋ค.
train.py also serves a live dashboard โ no web framework, just the stdlib http.server
(plus Pillow for JPEG encoding of frames): streamed render of the lander, telemetry HUD,
learning curve, ฮต schedule, and a hall of fame of the best landings.
ํ๊ณ / Limitations
- ์คํจ ์ฌ๋ก ๋๋ถ๋ถ์ ์ถ๋ฝ์ด ์๋๋ผ ์ฐฉ๋ฅ๋ ์์์ ๋งด๋๋ค 1000์คํ ํ์์์์ ๋๋ค. ฮต=0์ ๊ฒฐ์ ๋ก ์ ์ ์ฑ ์ด ๊ฐ์ง ์ ํ์ ์ฝ์ ์ ๋๋ค.
- ํ์คํธ์ฐจ๊ฐ ํฝ๋๋ค(ยฑ85). ์๋์ ๋ฐ๋ผ ํธ์ฐจ๊ฐ ์์ผ๋ ์ฌํ ์ ์ฐธ๊ณ ํ์ธ์.
- Downloads last month
- -
Evaluation results
- mean_reward on LunarLander-v3self-reported223.14 +/- 84.58
