DP-T pointflow: point-cloud diffusion policies for DexJoCo bimanual_assembly
Checkpoints of the pointflow arm of DP-T
(design and every result: docs/POINTFLOW_ARM.md in that repo). The policy is the decoupled
two-tower diffusion transformer of the released baseline
unikai/DP-T, with 128 unpooled object point tokens
(fixed CAD surface points posed at obj_pose) and 2 TCP tokens; nothing names a tip, a bore
or an axis, and no relation is computed for the policy.
| file | what | success, seeds 0-3 x 75 | fresh seeds 4-7 x 75 |
|---|---|---|---|
dpt_pointflow_gated_latest.ckpt (119 MB) |
gated cross-attention onto the point tokens, gates at zero, fine-tuned 100 epochs from the baseline's EMA weights | 55.0 % (baseline 33.7 %, p < 0.001) | 52.0 % (baseline 38.7 %, p = 0.0015) |
dpt_pointflow_tokens_flow_ep75.ckpt (71 MB) |
tokens in the decoder memory + pointflow auxiliary loss, trained from scratch, epoch 75 | 45.0 % | 47.7 % (baseline 38.7 %, p = 0.028) |
Paired closed-loop evaluation, eval_dp.py --reference-protocol, the default
rand_obj/bimanual_assembly eval config; per-episode records in outputs/eval/pointflow_*.
hf download b12902136/DP-T-pointflow dpt_pointflow_gated_latest.ckpt --local-dir outputs/pretrained
python eval_dp.py --checkpoint outputs/pretrained/dpt_pointflow_gated_latest.ckpt \
--eval-config $DEXJOCO_ROOT/configs/rand_obj/bimanual_assembly.yaml \
--episodes 75 --seed 0 --render-first 0 --no-render --reference-protocol
Inputs at inference: proprioception (state, 46) and the live object poses (obj_pose, 14);
the CAD point clouds are built inside the policy from the DexJoCo object XMLs. Like the
baseline, these are torch.save workspace payloads (cfg + model + EMA weights; the optimizer
state is dropped), loaded by eval_dp.py with
dexjoco_dp.policy.decoupled_pointflow_policy.DecoupledPointflowPolicy from the repo.