Walker2d Reflective Adaptation Policy (v1)
Model policy bipedal locomotion adaptif berbasis Critic-Guided Action Reflection untuk mengatasi actuator failure dan long-horizon numerical drift pada environment Walker2d.
Hasil Benchmark (10.000 Steps Evaluation)
- Normal Endurance Streak: 7.432 steps (+253.44m) vs Baseline Murni (4.901 steps).
- Fault Survival (Right Knee 100% Jammed): Bertahan 143 steps vs Baseline Murni (67 steps) — peningkatan +113%.
- Reflective Learning Rate (Alpha):
2.0e-03