Walker2d Calibrated Autonomous Reflective Agent (v2.1)
Model policy adaptif bipedal locomotion terkalibrasi yang menggabungkan:
- Intrinsically Robust Feedforward Actor (hasil konsolidasi LR $3.0\times 10^{-5}$).
- Filtered World Dynamics Anomaly Detector (Threshold: 8.0, Persistence: 3 frames).
- Online Critic-Guided Action Reflection Engine ($lpha = 2.0\times 10^{-3}$, 8 iterations).
Benchmark Highlights (1.000 Steps Suite)
- External Wind Gust (+140N Push): 1.000 steps penuh (+26.14 m) dengan hanya 5 trigger presisi (vs Baseline jatuh di step 367).
- Right Knee Jammed (0 Nm): Tembus 515 steps murni feedforward tanpa false alarm saat fase lari normal.