sim-square-broad-r01-mulligan-divl
Part of Mulligan. Browse evaluations at Policy Arena; the datasets are in the mulligan organization.
State-based agent with the parent's frozen diffusion actor and a distributional DIVL critic (policy.pt, stats.json).
| Property | Value |
|---|---|
| Task | sim-square-broad |
| Model round | R1 |
| Arm | mulligan |
| Campaign cell | sq_d1_r1_ours_mining_freecf_human_only |
| Seeds | 1, 2, 3, 4, 5 (one folder per seed) |
| Training step | 250001 |
Checkpoints
| Folder | Training run config |
|---|---|
seed-1 |
release/run-configs/sim-square-broad-r01-mulligan-divl__seed-1.json |
seed-2 |
release/run-configs/sim-square-broad-r01-mulligan-divl__seed-2.json |
seed-3 |
release/run-configs/sim-square-broad-r01-mulligan-divl__seed-3.json |
seed-4 |
release/run-configs/sim-square-broad-r01-mulligan-divl__seed-4.json |
seed-5 |
release/run-configs/sim-square-broad-r01-mulligan-divl__seed-5.json |
The frozen actor comes from sim-square-broad-r01-mulligan-idql.
Evaluations
Held-out initial-state grid; per-rollout outcomes are in the linked dataset.
| Dataset | Folder | N | Successes |
|---|---|---|---|
| sim-square-broad-r00-r03-eval | seed-1 |
32 | 21271/30000 |
| sim-square-broad-r00-r03-eval | seed-2 |
32 | 22553/30000 |
| sim-square-broad-r00-r03-eval | seed-3 |
32 | 21085/30000 |
| sim-square-broad-r00-r03-eval | seed-4 |
32 | 21323/30000 |
| sim-square-broad-r00-r03-eval | seed-5 |
32 | 21384/30000 |
Data
Trained on: sim-square-broad-c00-teleop-sobol, sim-square-broad-c01-dagger-mulligan, sim-square-broad-c01-sobol-policy-rollouts.
Referenced by the metadata of: sim-square-broad-r00-r03-eval.
Provenance
SHA-256 of every file is recorded in release.json. .pt files are PyTorch pickles; load them only in a trusted environment.
The run configs are in the Mulligan code repository; the recipes there retrain each checkpoint. License: mit.