Robotics
mulligan
sim-square-broad
divl-agent

sim-square-broad-r01-mulligan-divl

Part of Mulligan. Browse evaluations at Policy Arena; the datasets are in the mulligan organization.

State-based agent with the parent's frozen diffusion actor and a distributional DIVL critic (policy.pt, stats.json).

Property Value
Task sim-square-broad
Model round R1
Arm mulligan
Campaign cell sq_d1_r1_ours_mining_freecf_human_only
Seeds 1, 2, 3, 4, 5 (one folder per seed)
Training step 250001

Checkpoints

Folder Training run config
seed-1 release/run-configs/sim-square-broad-r01-mulligan-divl__seed-1.json
seed-2 release/run-configs/sim-square-broad-r01-mulligan-divl__seed-2.json
seed-3 release/run-configs/sim-square-broad-r01-mulligan-divl__seed-3.json
seed-4 release/run-configs/sim-square-broad-r01-mulligan-divl__seed-4.json
seed-5 release/run-configs/sim-square-broad-r01-mulligan-divl__seed-5.json

The frozen actor comes from sim-square-broad-r01-mulligan-idql.

Evaluations

Held-out initial-state grid; per-rollout outcomes are in the linked dataset.

Dataset Folder N Successes
sim-square-broad-r00-r03-eval seed-1 32 21271/30000
sim-square-broad-r00-r03-eval seed-2 32 22553/30000
sim-square-broad-r00-r03-eval seed-3 32 21085/30000
sim-square-broad-r00-r03-eval seed-4 32 21323/30000
sim-square-broad-r00-r03-eval seed-5 32 21384/30000

Data

Trained on: sim-square-broad-c00-teleop-sobol, sim-square-broad-c01-dagger-mulligan, sim-square-broad-c01-sobol-policy-rollouts.

Referenced by the metadata of: sim-square-broad-r00-r03-eval.

Provenance

SHA-256 of every file is recorded in release.json. .pt files are PyTorch pickles; load them only in a trusted environment. The run configs are in the Mulligan code repository; the recipes there retrain each checkpoint. License: mit.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Datasets used to train mulligan/sim-square-broad-r01-mulligan-divl

Collection including mulligan/sim-square-broad-r01-mulligan-divl