Robotics
Asimov
ONNX

Asimov Flex: base walk v0.8

Velocity-commanded walking policy for Asimov 1, trained in Isaac Lab with PPO. It is a base walk meant for later skills to build on: forward, backward and sideways walking, turning while walking or on the spot, and standing still.

Version 0.8 is better on stairs. It is walk v0.7 trained further with more practice on rough ground: at the end of every training episode, 40% of the robots on rough ground were put on the four hardest levels (±5.8 to ±9 cm). The rest of v0.7's training is unchanged. The goal was rough ground of ±8 cm; that is still not solved, but 10 cm stairs got much better. Training longer on very rough ground wore down the backward recovery step after a push from the front, so v0.8 is an earlier checkpoint (iteration 12450). The policy itself never sees the ground: it uses only the robot's joint and body sensors, so no camera or lidar is needed. The inputs, the outputs, the gait clock and the timing are the same as in v0.7. Earlier training for top speed, rough ground and stairs, for balance (pushes from every direction, steady forces), for differences between simulation and hardware (IMU tilt bias, motor strength and gain changes, extra mass, joint friction) and for the robot's timing (about 35 ms command delay at 50 Hz) is kept.

Tested in simulation only (Isaac Lab, and MuJoCo with the public Asimov 1 model). It has not been run on the real robot.

  • policy.onnx is the exported policy for inference, producing joint-position actions.
  • agent.yaml records the actor-critic and AMP training settings, including the motion data configuration.
  • env.yaml records the simulation and locomotion task settings, including observations, actions, commands, and rewards.
  • Training and evaluation code: menloresearch/isaac_asimov.

Inputs and outputs

  • Control rate 50 Hz. Actions: 23 joint-position offsets, target = default_pose + 0.25 * action, no filter.
  • 80 inputs: the same 78 as the standard Asimov 1 walking policy, in the same order, followed by a 2-value gait clock.
  • Gait clock: (sin(2 pi phi), cos(2 pi phi)), with phi in [0, 1).
    • Each policy step while walking: phi += elapsed / period, where elapsed is the measured real time since the previous policy step (0.02 s at 50 Hz). Use the measured time, not a fixed 0.02 s, if the loop runs slower.
    • period is 0.8 s at command speeds up to 0.2 m/s and 0.7 s at 0.8 m/s and above, linear in between (planar command speed).
    • When |v_xy| + |w_z| <= 0.1 the robot stands: the clock outputs (0, 0) and phi stays at 0.
    • Walking starts from phi = 0 (both feet down); the right foot lifts first.

Results in simulation

Measured at the robot's timing (50 Hz, 35 ms command delay). Terrain results in Isaac are from a new random seed with both versions on the same seed, 192 robots per case; the pooled rows add three more seeds (768 robots). MuJoCo uses 128 seeds per case. Falls are the share of test episodes.

Test v0.8 v0.7
Stairs up, 8 / 10 / 12 cm steps (Isaac) 1.6% / 3.1% / 88.0% falls 1.0% / 13.0% / 97.4%
10 cm steps up on 0.35 m treads (Isaac) 6.2% falls 31.2%
Rough ground, ±6 / ±8 cm (Isaac) 4.2% / 94.8% falls 7.3% / 96.4%
Four seeds pooled, 10 cm stairs up / rough ±6 cm / rough ±8 cm (Isaac) 3.6% / 9.9% / 73.8% falls 13.9% / 11.5% / 79.2%
Rough ground, ±6 cm (MuJoCo) 18.8% falls 26.6%
Stairs up 10 cm (MuJoCo) 17.2% falls 56.2%
Rough ground, ±2 / ±4 cm (MuJoCo) 2.3% / 7.0% falls 0% / 2.3%
Standing push 0.6 / 0.8 m/s (MuJoCo) 0% / 0% falls 0% / 0%
Falls on flat ground, all walking tests (256 robots each) 0% 0%
Forward speed at a 0.8 m/s command (Isaac) 0.718 m/s 0.702 m/s
Action jitter, mean of 11 tests 0.085 0.093
Energy per metre (cost of transport) at 0.25 / 0.5 / 0.8 m/s (Isaac) 0.649 / 0.682 / 0.730 0.732 / 0.667 / 0.684

Starting to walk and turn from a standstill: no falls in 1,024 robots. Every robot crossed a 10° slope going up. A 0.8 m/s push while walking (MuJoCo): 0.8% falls.

Known limits:

  • Rough ground with ±8 cm bumps is still not solved.
  • 12 cm steps going up still fail most of the time.
  • In MuJoCo, small rough ground (±2 / ±4 cm) has slightly more falls than with v0.7.
  • It uses more energy per metre at the top speed (+7% at 0.8 m/s, +2% at 0.5 m/s; 11% less at 0.25 m/s).
  • Slopes up are reliable only up to 10°; the ankle range limits steeper slopes.
  • With a large IMU tilt error (above about 7°), walking falls sooner than with v0.7 (at 10°, 62% against 0%). Keep the tilt error under 3°.
  • It has not been run on the real robot yet.
Downloads last month
64
Video Preview
loading