Instructions to use geonmin-kim/Cosmos3-Edge-Policy-SO101-init with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use geonmin-kim/Cosmos3-Edge-Policy-SO101-init with LeRobot:
- Notebooks
- Google Colab
- Kaggle
Cosmos3-Edge-Policy-SO101-init
nvidia/Cosmos3-Edge-Policy-DROID (4B Mixture-of-Transformers world action model,
DROID 8-D joint policy) converted to the LeRobot cosmos3 policy format and
retargeted to the SO-101 embodiment (5 joints + gripper, absolute joint position).
This is the un-finetuned initialization checkpoint: weights are bit-identical to
the NVIDIA release; only the config/normalization contract changed. No weight surgery
is needed because the Cosmos3 action head is a DomainAwareLinear over a zero-padded
64-dim action vector with a per-sample raw-dim mask.
Converted with scripts/convert_cosmos3_edge_policy_to_lerobot.py from
nota-github/xpu-lerobot#6.
SO-101 retargeting (vs the DROID contract)
| field | DROID | this repo |
|---|---|---|
raw_action_dim |
8 (7 joints + gripper) | 6 (5 joints + gripper) |
| cameras | wrist + left/right over-shoulder | wrist + top (3rd view zero-padded) |
| normalization | — | QUANTILES (q01/q99 → [-1,1]), stats injected from the training dataset at train time |
conditioning_fps |
15 | 30 |
invert_gripper |
true | false |
chunk_size / n_action_steps |
32 / 32 | 32 / 32 |
domain_id |
8 (droid_lerobot) |
8 (reused, adapted by fine-tuning) |
Usage
Requires the cosmos3 policy from the PR above and diffusers git main (>= 0.40.0.dev0
schema — 0.39 silently drops use_und_k_norm_for_gen/qk_norm_for_text/hidden_act
and builds a different architecture), huggingface_hub>=1.23, safetensors>=0.8.
lerobot-train \
--policy.path=geonmin-kim/Cosmos3-Edge-Policy-SO101-init \
--dataset.repo_id=<your SO-101 LeRobot dataset> ...
Tokenizer resolution is fully portable: text_processor_name_or_path and the
serialized processor's processor_name point at this repo id (with
local_files_only=false), and preprocessor_config.json ships so
Qwen3VLProcessor loads without the AutoTokenizer fallback. No overrides needed
on a fresh machine.
Fine-tuned descendants
- Flow matching (4-step):
geonmin-kim/Cosmos3-Edge-SO101-Morig-D6231-step{3000..30000}— useguidance_scale=1.0at inference (no prompt dropout during fine-tuning, CFG hurts). - Drift 1-step (
num_inference_steps=1):geonmin-kim/Cosmos3-Edge-SO101-Drift-D6231-step{5000..100000}— recommended for deployment (open-loop MAE 5.45 vs 26.96, latency 0.25s vs 0.62s per chunk on B200).
Training data of the descendants: geonmin-kim/SO101-large-small-cube-11dataset-ratio6231-0729.
- Downloads last month
- 21
Model tree for geonmin-kim/Cosmos3-Edge-Policy-SO101-init
Base model
nvidia/Cosmos3-Edge-Policy-DROID