mulligan/sim-square-narrow-c00-teleop-baseline
Franka Panda • Updated • 100 episodes • 254
Part of Mulligan. Browse evaluations at Policy Arena; the datasets are in the mulligan organization.
State-based agent with the parent's frozen diffusion actor and a distributional DIVL critic (policy.pt, stats.json).
| Property | Value |
|---|---|
| Task | sim-square-narrow |
| Model round | R0 |
| Arm | baseline |
| Campaign cell | sq_d0_r0_baseline_uniform |
| Seeds | 1, 2, 3, 4, 5 (one folder per seed) |
| Training step | 150001 |
| Folder | Training run config |
|---|---|
seed-1 |
release/run-configs/sim-square-narrow-r00-baseline-divl__seed-1.json |
seed-2 |
release/run-configs/sim-square-narrow-r00-baseline-divl__seed-2.json |
seed-3 |
release/run-configs/sim-square-narrow-r00-baseline-divl__seed-3.json |
seed-4 |
release/run-configs/sim-square-narrow-r00-baseline-divl__seed-4.json |
seed-5 |
release/run-configs/sim-square-narrow-r00-baseline-divl__seed-5.json |
The frozen actor comes from sim-square-narrow-r00-baseline-idql.
Held-out initial-state grid; per-rollout outcomes are in the linked dataset.
| Dataset | Folder | N | Successes |
|---|---|---|---|
| sim-square-narrow-r00-r03-eval | seed-1 |
32 | 5509/8000 |
| sim-square-narrow-r00-r03-eval | seed-2 |
32 | 5727/8000 |
| sim-square-narrow-r00-r03-eval | seed-3 |
32 | 5860/8000 |
| sim-square-narrow-r00-r03-eval | seed-4 |
32 | 5768/8000 |
| sim-square-narrow-r00-r03-eval | seed-5 |
32 | 5455/8000 |
Trained on: sim-square-narrow-c00-teleop-baseline.
Referenced by the metadata of: sim-square-narrow-r00-r03-eval.
SHA-256 of every file is recorded in release.json. .pt files are PyTorch pickles; load them only in a trusted environment.
The run configs are in the Mulligan code repository; the recipes there retrain each checkpoint. License: mit.