Square insertion flow specialist (MimicGen Square D1)
A small flow-matching policy for the robosuite / MimicGen Square task (put the square nut on the peg), trained by behaviour cloning on 44 demonstrations. It predicts 15-step chunks of Franka Panda joint-position offsets and a gripper command from two camera images and the robot state.
We use it as the demonstration-trained component (the "specialist" / task proxy) in our work on steering VLM-planned sampling policies with Proxy Policy Steering (PPS). It is not a strong standalone policy: alone it solves 5 of our 50 test scenes (see Evaluation).
Built with DINOv3. The image encoder is Meta's DINOv3 ViT-S/16
(facebook/dinov3-vits16-pretrain-lvd1689m), fine-tuned end to end. These weights are therefore a
derivative of the DINO Materials and are distributed under the DINOv3 License (LICENSE.md).
Model
| Architecture | ProxyPytorch (flow variant): DINOv3 ViT-S/16 image encoder + 8-layer Gemma-style action expert (gemma_12m: width 384, MLP 512, 8 heads, 1 KV head, head dim 128) + linear state/action/time projections |
| Parameters | 33.9 M (21.6 M in the DINOv3 encoder), float32 |
| Attention | two_block_diffusion: image and state tokens form one block, the noisy action chunk the second |
| Language | The openpi input pipeline tokenizes the prompt, but this model does not attend to language tokens |
| Weights | model.safetensors = EMA weights (decay 0.999) at step 20,000 |
Inputs
- Images: table (agentview) and wrist (eye-in-hand) cameras, RGB, 224 × 224. The demonstrations were re-rendered at 224 px from the stored MuJoCo states with visual geometry only.
- State (8): the 7 arm joint positions (rad) and the gripper closure
clip((0.080 - aperture) / 0.080, 0, 1)(1 = closed), with aperture = the finger width in metres (robot0_gripper_qpos[0] - robot0_gripper_qpos[1]).
Outputs
An action chunk of shape 15 × 8 (one row per control step):
- Columns 0–6: joint-position targets as offsets from the current joints (rad), normalised per
row and column with
mean/stdinaction_norm_stats.json(action_norm = demo_delta, pooled over the demonstrations). Target joints = current joints + (a · std + mean). - Column 7: gripper command, 0 = open and 1 = closed, normalised as (g − 0.5) / 0.5.
Flow convention: x_t = (1 - t) · a + t · ε, and the network predicts the velocity ε - a.
Sampling integrates from t = 1 (noise) to t = 0 with 10 Euler steps (sample_actions(num_steps=10)).
Training data
- MimicGen Square D1 (
Square_D1, MimicGen's variant of robosuiteNutAssemblySquarewith a wider initial-state distribution than D0; Franka Panda): a split of 49 generated demonstrations (demo_0–demo_49withoutdemo_22), 44 for training and 5 held out (demo_45–demo_49). The demonstration names are listed inbc_metadata.json. - Actions relabelled for a joint-position controller: the target at step t is the joint position at step t + 1, plus the gripper command mapped from −1 / +1 to 0 / 1.
- 6,032 training windows (stride 1).
Training
| Config | flow_task_square_bc (openpi fork, openpi/src/openpi/training/config.py) |
| Steps / batch | 20,000 / 32 |
| Optimiser | AdamW, peak learning rate 2.5e-5, 1,000 warm-up steps, cosine decay to 2.5e-6 |
| Flow time | beta(1.5, 1) · 0.999 + 0.001 |
| Augmentation | random image shift of up to 4 px |
| EMA | 0.999 |
| Encoder | DINOv3 not frozen |
All settings are also in bc_metadata.json.
Evaluation
On our 50-scene Square D1 test panel (robosuite Square_D1, the same task variant as the training data, Panda with a
joint-position controller, scenes 1–50), the policy alone succeeds in 5 of 50 (10%). Success is
the task's own check: the nut rests on the peg.
In our steering experiments it is combined with a VLM-planned sampling-based base policy. The policy is only useful as a steering component, not as a standalone controller.
Files
| File | Content |
|---|---|
model.safetensors |
EMA weights, step 20,000 |
action_norm_stats.json |
per-row action mean / std (15 × 8) and the normalisation settings |
bc_metadata.json |
training settings and the demonstration split |
metadata.pt |
checkpoint step and time stamp |
assets/local/isaaclab_capsule_score/norm_stats.json |
data-pipeline normalisation file required by the openpi checkpoint layout |
LICENSE.md |
the DINOv3 License |
Usage
With the openpi fork that defines ProxyPytorch and the config flow_task_square_bc:
import safetensors.torch
import openpi.models_pytorch.proxy_pytorch as proxy
import openpi.training.config as config
cfg = config.get_config("flow_task_square_bc")
model = proxy.ProxyPytorch(cfg.model).eval()
safetensors.torch.load_model(model, "model.safetensors")
# model.sample_actions(device, observation, num_steps=10) -> normalised 15 x 8 chunk
Build the observation with the same input transform as training (two images, image masks, the
8-dimensional state). Then de-normalise the chunk with action_norm_stats.json as described in
Outputs.
Limitations
- Simulation only (robosuite / MimicGen Square D1), one robot (Franka Panda), one camera set-up.
- Trained on 44 demonstrations; standalone success is low (10%).
- The actions are joint-position offsets for a joint-position controller. They do not transfer directly to other controllers or robots.
License
The weights contain a fine-tuned copy of DINOv3 and are released under the DINOv3 License
(LICENSE.md, from facebookresearch/dinov3).
Any use or redistribution must follow its terms, including the trade-control and acceptable-use
conditions. If you publish research that uses this model, acknowledge the use of DINOv3, as the
license requires.
Model tree for jsiburian/square-insertion-flow-specialist
Base model
facebook/dinov3-vit7b16-pretrain-lvd1689m