Robotics
mulligan
sim-square-narrow
divl-agent

sim-square-narrow-r00-baseline-divl

Part of Mulligan. Browse evaluations at Policy Arena; the datasets are in the mulligan organization.

State-based agent with the parent's frozen diffusion actor and a distributional DIVL critic (policy.pt, stats.json).

Property Value
Task sim-square-narrow
Model round R0
Arm baseline
Campaign cell sq_d0_r0_baseline_uniform
Seeds 1, 2, 3, 4, 5 (one folder per seed)
Training step 150001

Checkpoints

Folder Training run config
seed-1 release/run-configs/sim-square-narrow-r00-baseline-divl__seed-1.json
seed-2 release/run-configs/sim-square-narrow-r00-baseline-divl__seed-2.json
seed-3 release/run-configs/sim-square-narrow-r00-baseline-divl__seed-3.json
seed-4 release/run-configs/sim-square-narrow-r00-baseline-divl__seed-4.json
seed-5 release/run-configs/sim-square-narrow-r00-baseline-divl__seed-5.json

The frozen actor comes from sim-square-narrow-r00-baseline-idql.

Evaluations

Held-out initial-state grid; per-rollout outcomes are in the linked dataset.

Dataset Folder N Successes
sim-square-narrow-r00-r03-eval seed-1 32 5509/8000
sim-square-narrow-r00-r03-eval seed-2 32 5727/8000
sim-square-narrow-r00-r03-eval seed-3 32 5860/8000
sim-square-narrow-r00-r03-eval seed-4 32 5768/8000
sim-square-narrow-r00-r03-eval seed-5 32 5455/8000

Data

Trained on: sim-square-narrow-c00-teleop-baseline.

Referenced by the metadata of: sim-square-narrow-r00-r03-eval.

Provenance

SHA-256 of every file is recorded in release.json. .pt files are PyTorch pickles; load them only in a trusted environment. The run configs are in the Mulligan code repository; the recipes there retrain each checkpoint. License: mit.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Dataset used to train mulligan/sim-square-narrow-r00-baseline-divl

Collection including mulligan/sim-square-narrow-r00-baseline-divl