Square insertion flow specialist (MimicGen Square D1)

A small flow-matching policy for the robosuite / MimicGen Square task (put the square nut on the peg), trained by behaviour cloning on 44 demonstrations. It predicts 15-step chunks of Franka Panda joint-position offsets and a gripper command from two camera images and the robot state.

We use it as the demonstration-trained component (the "specialist" / task proxy) in our work on steering VLM-planned sampling policies with Proxy Policy Steering (PPS). It is not a strong standalone policy: alone it solves 5 of our 50 test scenes (see Evaluation).

Built with DINOv3. The image encoder is Meta's DINOv3 ViT-S/16 (facebook/dinov3-vits16-pretrain-lvd1689m), fine-tuned end to end. These weights are therefore a derivative of the DINO Materials and are distributed under the DINOv3 License (LICENSE.md).

Model

Architecture ProxyPytorch (flow variant): DINOv3 ViT-S/16 image encoder + 8-layer Gemma-style action expert (gemma_12m: width 384, MLP 512, 8 heads, 1 KV head, head dim 128) + linear state/action/time projections
Parameters 33.9 M (21.6 M in the DINOv3 encoder), float32
Attention two_block_diffusion: image and state tokens form one block, the noisy action chunk the second
Language The openpi input pipeline tokenizes the prompt, but this model does not attend to language tokens
Weights model.safetensors = EMA weights (decay 0.999) at step 20,000

Inputs

  • Images: table (agentview) and wrist (eye-in-hand) cameras, RGB, 224 × 224. The demonstrations were re-rendered at 224 px from the stored MuJoCo states with visual geometry only.
  • State (8): the 7 arm joint positions (rad) and the gripper closure clip((0.080 - aperture) / 0.080, 0, 1) (1 = closed), with aperture = the finger width in metres (robot0_gripper_qpos[0] - robot0_gripper_qpos[1]).

Outputs

An action chunk of shape 15 × 8 (one row per control step):

  • Columns 0–6: joint-position targets as offsets from the current joints (rad), normalised per row and column with mean / std in action_norm_stats.json (action_norm = demo_delta, pooled over the demonstrations). Target joints = current joints + (a · std + mean).
  • Column 7: gripper command, 0 = open and 1 = closed, normalised as (g − 0.5) / 0.5.

Flow convention: x_t = (1 - t) · a + t · ε, and the network predicts the velocity ε - a. Sampling integrates from t = 1 (noise) to t = 0 with 10 Euler steps (sample_actions(num_steps=10)).

Training data

  • MimicGen Square D1 (Square_D1, MimicGen's variant of robosuite NutAssemblySquare with a wider initial-state distribution than D0; Franka Panda): a split of 49 generated demonstrations (demo_0–demo_49 without demo_22), 44 for training and 5 held out (demo_45–demo_49). The demonstration names are listed in bc_metadata.json.
  • Actions relabelled for a joint-position controller: the target at step t is the joint position at step t + 1, plus the gripper command mapped from −1 / +1 to 0 / 1.
  • 6,032 training windows (stride 1).

Training

Config flow_task_square_bc (openpi fork, openpi/src/openpi/training/config.py)
Steps / batch 20,000 / 32
Optimiser AdamW, peak learning rate 2.5e-5, 1,000 warm-up steps, cosine decay to 2.5e-6
Flow time beta(1.5, 1) · 0.999 + 0.001
Augmentation random image shift of up to 4 px
EMA 0.999
Encoder DINOv3 not frozen

All settings are also in bc_metadata.json.

Evaluation

On our 50-scene Square D1 test panel (robosuite Square_D1, the same task variant as the training data, Panda with a joint-position controller, scenes 1–50), the policy alone succeeds in 5 of 50 (10%). Success is the task's own check: the nut rests on the peg.

In our steering experiments it is combined with a VLM-planned sampling-based base policy. The policy is only useful as a steering component, not as a standalone controller.

Files

File Content
model.safetensors EMA weights, step 20,000
action_norm_stats.json per-row action mean / std (15 × 8) and the normalisation settings
bc_metadata.json training settings and the demonstration split
metadata.pt checkpoint step and time stamp
assets/local/isaaclab_capsule_score/norm_stats.json data-pipeline normalisation file required by the openpi checkpoint layout
LICENSE.md the DINOv3 License

Usage

With the openpi fork that defines ProxyPytorch and the config flow_task_square_bc:

import safetensors.torch
import openpi.models_pytorch.proxy_pytorch as proxy
import openpi.training.config as config

cfg = config.get_config("flow_task_square_bc")
model = proxy.ProxyPytorch(cfg.model).eval()
safetensors.torch.load_model(model, "model.safetensors")
# model.sample_actions(device, observation, num_steps=10) -> normalised 15 x 8 chunk

Build the observation with the same input transform as training (two images, image masks, the 8-dimensional state). Then de-normalise the chunk with action_norm_stats.json as described in Outputs.

Limitations

  • Simulation only (robosuite / MimicGen Square D1), one robot (Franka Panda), one camera set-up.
  • Trained on 44 demonstrations; standalone success is low (10%).
  • The actions are joint-position offsets for a joint-position controller. They do not transfer directly to other controllers or robots.

License

The weights contain a fine-tuned copy of DINOv3 and are released under the DINOv3 License (LICENSE.md, from facebookresearch/dinov3). Any use or redistribution must follow its terms, including the trade-control and acceptable-use conditions. If you publish research that uses this model, acknowledge the use of DINOv3, as the license requires.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
33.9M params
Tensor type
F32
·
Video Preview
loading

Model tree for jsiburian/square-insertion-flow-specialist