strands-isaaclab-shadow-handover-policy

rsl_rl PPO policy for two Shadow hands handing an object over (Isaac-Shadow-Handover), trained with strands-robots' isaaclab train_policy provider (PR #4227). Rollouts recorded through strands: cagataydev/strands-isaaclab-shadow-handover.

playback: 4 recorded envs

Results

2048 envs × 1500 PPO iterations (49 M env steps, 45 min 35 s, Newton/MJWarp, 1× L40S): mean reward 0.00 → 23.3, handover success 0.89. Recorded 12 rollouts: mean return 13.0.

PPO iteration mean reward handover success env-steps/s
0 0.00 0.000 27,700
100 0.32 0.002 19,498
300 15.47 0.638 40,439
600 19.51 0.767 16,978
900 20.58 0.791 15,199
1200 24.09 0.939 14,530
1499 23.26 0.892 14,438

Files

file what
model_1499.pt rsl_rl checkpoint (actor-critic + optimizer), final iteration
exported/policy.pt TorchScript actor incl. observation normalizer; input (N, 290) → (N, 40) (20 actuated joints per hand)
exported/policy.onnx (+ .data) ONNX export of the same actor
params/agent.yaml, params/env.yaml exact Isaac Lab agent / env configs of this run
train_curve.json per-iteration training log parsed by the strands provider
record.json, verify.json recording + verification reports
playback.gif, playback.mp4, frame.png rendered rollouts
examples/ recording script used to make the dataset

How it was made with strands-robots

# Isaac Lab in its OWN venv (its pins clash with strands; strands never imports it)
uv venv --python 3.12 ~/il && uv pip install --python ~/il/bin/python --prerelease=allow \
  --index https://pypi.nvidia.com --index-strategy unsafe-best-match "isaaclab[rsl-rl,isaacsim]==3.0.0rc1"
export ISAACLAB_PYTHON=~/il/bin/python
export OMNI_KIT_ACCEPT_EULA=YES          # you accept the NVIDIA Omniverse / Isaac Sim EULA yourself
pip install "git+https://github.com/cagataycali/robots@feat/isaaclab-trainer"   # strands-robots with PR #4227

As an agent tool call (the train_policy tool is a Strands @tool):

from strands import Agent
from strands_robots.tools.train_policy import train_policy

agent = Agent(tools=[train_policy])
agent("Train two Shadow hands to hand an object over with the isaaclab provider: task Isaac-Shadow-Handover, 1500 iterations, seed 1.")
# -> train_policy(action="train", provider="isaaclab", steps=1500, seed=1, output_dir="runs/c6_shadow_handover",
#                 extra={"task": "Isaac-Shadow-Handover", "timeout_s": 10800})

As plain Python (exactly what produced this run):

from strands_robots.tools.train_policy import train_policy

job = train_policy(action="train", provider="isaaclab", steps=1500, seed=1, output_dir="runs/c6_shadow_handover",
                   extra={"task": "Isaac-Shadow-Handover", "timeout_s": 10800})   # task default: 2048 envs, Newton/MJWarp physics
train_policy(action="status", provider="isaaclab", job_id="<job_id from the result>")

Under the hood: python -m isaaclab train --rl_library rsl_rl --task Isaac-Shadow-Handover --max_iterations 1500 --seed 1 in $ISAACLAB_PYTHON. Job id of this run: isaaclab-20260929-081406-931b9b40e4ff. Docs: docs/learn/training/isaaclab.md · PR: strands-labs/robots#4227.

Use it

Play it in Isaac Lab:

huggingface-cli download cagataydev/strands-isaaclab-shadow-handover-policy --local-dir handover_policy
OMNI_KIT_ACCEPT_EULA=YES $ISAACLAB_PYTHON -m isaaclab play --rl_library rsl_rl --task Isaac-Shadow-Handover \
  --num_envs 16 --checkpoint handover_policy/model_1499.pt     # add --video --video_length 300

Or use the exported actor anywhere (TorchScript):

import torch
pi = torch.jit.load("handover_policy/exported/policy.pt").eval()
actions = pi(obs)          # obs: (N, 290) Isaac Lab 'policy' observation group -> (N, 40)

As a strands Policy (what the recording used):

import json, numpy as np, torch
from strands_robots.policies.base import Policy

class RslRlJitPolicy(Policy):
    def __init__(self, jit_path, action_names, device="cpu"):
        self.net = torch.jit.load(jit_path, map_location=device).eval()
        self.device, self.action_names, self.robot_state_keys = device, list(action_names), []
    @property
    def provider_name(self): return "isaaclab_rsl_rl_jit"
    def set_robot_state_keys(self, keys): self.robot_state_keys = list(keys)
    def reset(self, seed=None): self.net.reset()
    async def get_actions(self, observation_dict, instruction, **kw):
        x = torch.as_tensor(np.asarray(observation_dict["policy_obs"], np.float32), device=self.device)
        with torch.inference_mode():
            y = self.net(x.unsqueeze(0))[0].cpu().numpy()
        return [dict(zip(self.action_names, map(float, y)))]

names = json.load(open("handover_policy/record.json"))["action_names"]
policy = RslRlJitPolicy("handover_policy/exported/policy.pt", names)
chunk = policy.get_actions_sync({"policy_obs": obs_290}, "hand the object over")   # [{"right_hand.rh_WRJ2": ..., ..., "left_hand.lh_THJ1": ...}]  (40 = 2 x 20 actuated joints)

Record your own dataset with strands:

The final checkpoint was rolled out and recorded with examples/record_trained_policy.py (included): rebuild the task env in play mode with an RTX camera → load model_1499.pt with rsl_rl and export TorchScript/ONNX → wrap the actor as a strands Policy (RslRlJitPolicy, max |Δa| vs rsl_rl = 1.5e-06) → step with policy.get_actions_sync(...) → write every frame through strands DatasetRecorder (LeRobot v3) → verify with strands verify_dataset.

OMNI_KIT_ACCEPT_EULA=YES PYTHONPATH=/path/to/strands-robots $ISAACLAB_PYTHON examples/record_trained_policy.py \
  --task Isaac-Shadow-Handover --checkpoint model_1499.pt --episodes 12 --frames 300 \
  --cam fixed --cam_name front --eye 1.0,-1.5,1.05 --target 0,-0.5,0.55 \
  --task_str "pass the object from one hand to the other and bring it to the goal" --robot_type shadow_hand_x2 --root out/ds --repo_id cagataydev/strands-isaaclab-shadow-handover
$ISAACLAB_PYTHON examples/record_trained_policy.py --verify out/ds --repo_id cagataydev/strands-isaaclab-shadow-handover

Provenance

  • strands-robots: feat/isaaclab-trainer @ fa66fc68 — strands-labs/robots#4227 (isaaclab train_policy provider; DatasetRecorder; verify_dataset)
  • Isaac Lab 3.0.0rc1 · Isaac Sim 6.1.0.0 · Newton / MJWarp (task default physics) · rsl-rl-lib 5.4.1 (PPO) · lerobot 0.6.1
  • GPU: 1× NVIDIA L40S (46 GB), shared with the Franka cable-lift training
  • Seeds: training seed 1 (params/agent.yaml, params/env.yaml); recording seed 7
  • Training job: isaaclab-20260929-081406-931b9b40e4ff, 2026-09-29

Limitations

  • Simulation only; no real Shadow Hands were used; no sim-to-real claims.
  • Release candidates: Isaac Lab 3.0.0rc1 / Isaac Sim 6.1.0.0. Trained on the task's default Newton/MJWarp physics; the checkpoint does not remember the preset (IL-X-011) — replay on the same physics.
  • Not every rollout succeeds: training success is 0.89 (peak 0.96), and in the recording 3 of 12 episodes ended early when the object was dropped (lengths 43, 49 and 65 frames); the other 9 ran the full 5 s.
  • Recording: 12 parallel envs, 300 frames (5 s @ 60 fps) each from one fixed camera; an episode ends at 300 frames or when the object is dropped. observation.state = 24 joint pos of the right hand only + 290-D two-hand policy observation + root pos/quat (321-D); the left hand's state is inside policy_obs.*. Not a standard LeRobot robot state.
  • create_policy("rl") in strands cannot load rsl_rl checkpoints yet (IL-X-006); use the exported TorchScript + the wrapper above.

License

Card choice: license: other — our generated data / weights under CC-BY-4.0, plus NVIDIA notices. Why:

  • Trajectories, rendered camera video, playback clips and trained weights are user-generated content made with NVIDIA Isaac Sim / Isaac Lab; the NVIDIA Omniverse License Agreement (governs Isaac Sim 6.1, isaacsim/LICENSE.txt) §2.1 allows distributing "user generated content that you develop using Omniverse, such as video, audio, stills, models, 3D assets and screen captures". We release it under CC-BY-4.0.
  • No NVIDIA Content is redistributed: the Shadow Hand USDs (Robots_Multiphysics/ShadowRobot/ShadowHandMultiPhysics_v0/…) and the ground-plane asset come from the Isaac Lab / Isaac Sim asset packs and are not in this repo (params/env.yaml only references their paths). "Shadow Hand" is a product of The Shadow Robot Company; no endorsement by Shadow Robot or NVIDIA is implied.
  • params/*.yaml are Isaac Lab configurations (BSD-3-Clause); examples/* are Apache-2.0 like strands-robots. Running Isaac Sim requires your own acceptance of the NVIDIA Isaac Sim / Omniverse EULA. Full text: LICENSE.md.
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Dataset used to train cagataydev/strands-isaaclab-shadow-handover-policy

Evaluation results

  • handover success rate (last training iteration) on Isaac-Shadow-Handover (Isaac Lab 3.0.0rc1, 2048 envs)
    self-reported
    0.892
  • mean episode reward (last training iteration) on Isaac-Shadow-Handover (Isaac Lab 3.0.0rc1, 2048 envs)
    self-reported
    23.260