Instructions to use cagataydev/strands-isaaclab-shadow-handover-policy with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use cagataydev/strands-isaaclab-shadow-handover-policy with LeRobot:
- Notebooks
- Google Colab
- Kaggle
strands-isaaclab-shadow-handover-policy
rsl_rl PPO policy for two Shadow hands handing an object over (Isaac-Shadow-Handover), trained with strands-robots' isaaclab
train_policy provider (PR #4227). Rollouts recorded through strands: cagataydev/strands-isaaclab-shadow-handover.
Results
2048 envs × 1500 PPO iterations (49 M env steps, 45 min 35 s, Newton/MJWarp, 1× L40S): mean reward 0.00 → 23.3, handover success 0.89. Recorded 12 rollouts: mean return 13.0.
| PPO iteration | mean reward | handover success | env-steps/s |
|---|---|---|---|
| 0 | 0.00 | 0.000 | 27,700 |
| 100 | 0.32 | 0.002 | 19,498 |
| 300 | 15.47 | 0.638 | 40,439 |
| 600 | 19.51 | 0.767 | 16,978 |
| 900 | 20.58 | 0.791 | 15,199 |
| 1200 | 24.09 | 0.939 | 14,530 |
| 1499 | 23.26 | 0.892 | 14,438 |
Files
| file | what |
|---|---|
model_1499.pt |
rsl_rl checkpoint (actor-critic + optimizer), final iteration |
exported/policy.pt |
TorchScript actor incl. observation normalizer; input (N, 290) → (N, 40) (20 actuated joints per hand) |
exported/policy.onnx (+ .data) |
ONNX export of the same actor |
params/agent.yaml, params/env.yaml |
exact Isaac Lab agent / env configs of this run |
train_curve.json |
per-iteration training log parsed by the strands provider |
record.json, verify.json |
recording + verification reports |
playback.gif, playback.mp4, frame.png |
rendered rollouts |
examples/ |
recording script used to make the dataset |
How it was made with strands-robots
# Isaac Lab in its OWN venv (its pins clash with strands; strands never imports it)
uv venv --python 3.12 ~/il && uv pip install --python ~/il/bin/python --prerelease=allow \
--index https://pypi.nvidia.com --index-strategy unsafe-best-match "isaaclab[rsl-rl,isaacsim]==3.0.0rc1"
export ISAACLAB_PYTHON=~/il/bin/python
export OMNI_KIT_ACCEPT_EULA=YES # you accept the NVIDIA Omniverse / Isaac Sim EULA yourself
pip install "git+https://github.com/cagataycali/robots@feat/isaaclab-trainer" # strands-robots with PR #4227
As an agent tool call (the train_policy tool is a Strands @tool):
from strands import Agent
from strands_robots.tools.train_policy import train_policy
agent = Agent(tools=[train_policy])
agent("Train two Shadow hands to hand an object over with the isaaclab provider: task Isaac-Shadow-Handover, 1500 iterations, seed 1.")
# -> train_policy(action="train", provider="isaaclab", steps=1500, seed=1, output_dir="runs/c6_shadow_handover",
# extra={"task": "Isaac-Shadow-Handover", "timeout_s": 10800})
As plain Python (exactly what produced this run):
from strands_robots.tools.train_policy import train_policy
job = train_policy(action="train", provider="isaaclab", steps=1500, seed=1, output_dir="runs/c6_shadow_handover",
extra={"task": "Isaac-Shadow-Handover", "timeout_s": 10800}) # task default: 2048 envs, Newton/MJWarp physics
train_policy(action="status", provider="isaaclab", job_id="<job_id from the result>")
Under the hood: python -m isaaclab train --rl_library rsl_rl --task Isaac-Shadow-Handover --max_iterations 1500 --seed 1 in $ISAACLAB_PYTHON.
Job id of this run: isaaclab-20260929-081406-931b9b40e4ff. Docs: docs/learn/training/isaaclab.md · PR: strands-labs/robots#4227.
Use it
Play it in Isaac Lab:
huggingface-cli download cagataydev/strands-isaaclab-shadow-handover-policy --local-dir handover_policy
OMNI_KIT_ACCEPT_EULA=YES $ISAACLAB_PYTHON -m isaaclab play --rl_library rsl_rl --task Isaac-Shadow-Handover \
--num_envs 16 --checkpoint handover_policy/model_1499.pt # add --video --video_length 300
Or use the exported actor anywhere (TorchScript):
import torch
pi = torch.jit.load("handover_policy/exported/policy.pt").eval()
actions = pi(obs) # obs: (N, 290) Isaac Lab 'policy' observation group -> (N, 40)
As a strands Policy (what the recording used):
import json, numpy as np, torch
from strands_robots.policies.base import Policy
class RslRlJitPolicy(Policy):
def __init__(self, jit_path, action_names, device="cpu"):
self.net = torch.jit.load(jit_path, map_location=device).eval()
self.device, self.action_names, self.robot_state_keys = device, list(action_names), []
@property
def provider_name(self): return "isaaclab_rsl_rl_jit"
def set_robot_state_keys(self, keys): self.robot_state_keys = list(keys)
def reset(self, seed=None): self.net.reset()
async def get_actions(self, observation_dict, instruction, **kw):
x = torch.as_tensor(np.asarray(observation_dict["policy_obs"], np.float32), device=self.device)
with torch.inference_mode():
y = self.net(x.unsqueeze(0))[0].cpu().numpy()
return [dict(zip(self.action_names, map(float, y)))]
names = json.load(open("handover_policy/record.json"))["action_names"]
policy = RslRlJitPolicy("handover_policy/exported/policy.pt", names)
chunk = policy.get_actions_sync({"policy_obs": obs_290}, "hand the object over") # [{"right_hand.rh_WRJ2": ..., ..., "left_hand.lh_THJ1": ...}] (40 = 2 x 20 actuated joints)
Record your own dataset with strands:
The final checkpoint was rolled out and recorded with examples/record_trained_policy.py
(included): rebuild the task env in play mode with an RTX camera → load model_1499.pt with rsl_rl and export TorchScript/ONNX →
wrap the actor as a strands Policy (RslRlJitPolicy, max |Δa| vs rsl_rl = 1.5e-06) → step with
policy.get_actions_sync(...) → write every frame through strands DatasetRecorder (LeRobot v3) → verify with strands verify_dataset.
OMNI_KIT_ACCEPT_EULA=YES PYTHONPATH=/path/to/strands-robots $ISAACLAB_PYTHON examples/record_trained_policy.py \
--task Isaac-Shadow-Handover --checkpoint model_1499.pt --episodes 12 --frames 300 \
--cam fixed --cam_name front --eye 1.0,-1.5,1.05 --target 0,-0.5,0.55 \
--task_str "pass the object from one hand to the other and bring it to the goal" --robot_type shadow_hand_x2 --root out/ds --repo_id cagataydev/strands-isaaclab-shadow-handover
$ISAACLAB_PYTHON examples/record_trained_policy.py --verify out/ds --repo_id cagataydev/strands-isaaclab-shadow-handover
Provenance
- strands-robots:
feat/isaaclab-trainer@fa66fc68— strands-labs/robots#4227 (isaaclabtrain_policy provider;DatasetRecorder;verify_dataset) - Isaac Lab 3.0.0rc1 · Isaac Sim 6.1.0.0 · Newton / MJWarp (task default physics) · rsl-rl-lib 5.4.1 (PPO) · lerobot 0.6.1
- GPU: 1× NVIDIA L40S (46 GB), shared with the Franka cable-lift training
- Seeds: training seed 1 (
params/agent.yaml,params/env.yaml); recording seed 7 - Training job:
isaaclab-20260929-081406-931b9b40e4ff, 2026-09-29
Limitations
- Simulation only; no real Shadow Hands were used; no sim-to-real claims.
- Release candidates: Isaac Lab 3.0.0rc1 / Isaac Sim 6.1.0.0. Trained on the task's default Newton/MJWarp physics; the checkpoint does not remember the preset (IL-X-011) — replay on the same physics.
- Not every rollout succeeds: training success is 0.89 (peak 0.96), and in the recording 3 of 12 episodes ended early when the object was dropped (lengths 43, 49 and 65 frames); the other 9 ran the full 5 s.
- Recording: 12 parallel envs, 300 frames (5 s @ 60 fps) each from one fixed camera; an episode ends at 300 frames or when the
object is dropped.
observation.state= 24 joint pos of the right hand only + 290-D two-hand policy observation + root pos/quat (321-D); the left hand's state is insidepolicy_obs.*. Not a standard LeRobot robot state. create_policy("rl")in strands cannot load rsl_rl checkpoints yet (IL-X-006); use the exported TorchScript + the wrapper above.
License
Card choice: license: other — our generated data / weights under CC-BY-4.0, plus NVIDIA notices. Why:
- Trajectories, rendered camera video, playback clips and trained weights are user-generated content made with NVIDIA Isaac Sim /
Isaac Lab; the NVIDIA Omniverse License Agreement (governs Isaac Sim 6.1,
isaacsim/LICENSE.txt) §2.1 allows distributing "user generated content that you develop using Omniverse, such as video, audio, stills, models, 3D assets and screen captures". We release it under CC-BY-4.0. - No NVIDIA Content is redistributed: the Shadow Hand USDs (
Robots_Multiphysics/ShadowRobot/ShadowHandMultiPhysics_v0/…) and the ground-plane asset come from the Isaac Lab / Isaac Sim asset packs and are not in this repo (params/env.yamlonly references their paths). "Shadow Hand" is a product of The Shadow Robot Company; no endorsement by Shadow Robot or NVIDIA is implied. params/*.yamlare Isaac Lab configurations (BSD-3-Clause);examples/*are Apache-2.0 like strands-robots. Running Isaac Sim requires your own acceptance of the NVIDIA Isaac Sim / Omniverse EULA. Full text:LICENSE.md.
Dataset used to train cagataydev/strands-isaaclab-shadow-handover-policy
Evaluation results
- handover success rate (last training iteration) on Isaac-Shadow-Handover (Isaac Lab 3.0.0rc1, 2048 envs)self-reported0.892
- mean episode reward (last training iteration) on Isaac-Shadow-Handover (Isaac Lab 3.0.0rc1, 2048 envs)self-reported23.260
