DiVE β€” Learning Decomposed Visibility for Efficient Active Exploration of Cluttered Scenes

Checkpoints for DiVE, published at the Conference on Robot Learning (CoRL) 2026.

DiVE decomposes scene visibility into what a viewpoint change can reveal and what a push can reveal, and uses that decomposition to decide between looking and interacting before ranking candidate actions within the chosen mode.

Contents

Folder Paper notation Role
DiVE_decomposed_visibility_estimator/ f_DiVE Predicts the view-resolvable (phi_view) and push-resolvable (phi_push) visibility maps
view_selection_policy/ pi_view Ranks the 98 candidate viewpoints
push_selection_policy/ pi_push Ranks the candidate pushes
belief_map_estimator/ f_scene Recursive semantic belief updater used as the central tracker

The three DiVE folders each hold a single best.pth saved by the training script, with keys epoch, model_state_dict, optimizer_state_dict, scheduler_state_dict, scaler_state_dict and the validation metrics of that epoch. Load the weights with torch.load(path, map_location="cpu")["model_state_dict"]. belief_map_estimator/estimator.pt is a plain state_dict instead and loads directly.

The voxel grid is (D, H, W) = (60, 120, 80) throughout.


DiVE_decomposed_visibility_estimator β€” decomposed visibility estimator

  • Model: support_code/model_UAB_unified.py β†’ UNet3DPushUnified
  • Trainer: 2_belief_unified/train_UAB_unified.py
  • Architecture: 3D U-Net with a shared backbone and a 2-channel output head, 0.34M parameters
  • Input: (B, 3, 60, 120, 80) = [phi_view_t, phi_push_t, voxel_t]
  • Output: (B, 2, 60, 120, 80) raw logits; apply sigmoid for visibility in [0, 1]
  • Trained for 200 epochs with teacher-forcing annealing; the released checkpoint is epoch 199 (self-feeding validation loss 3.92 β€” view head 2.48, push head 1.44)
  • Original run name: UAB_unified_ycb_v4_ep200_TFanneal

view_selection_policy β€” viewpoint selection policy

  • Model: support_code/model_reward.py β†’ RewardNet
  • Trainer: 4_policy/train_reward_UA.py
  • Architecture: 3D conv encoder collapsed onto the 7x14 camera grid, then a 2D conv head (base_ch=16, use_seen_mask=True), 1.02M parameters
  • Input: belief map (B, 1, 60, 120, 80) plus a (B, 7, 14) binary mask of the viewpoints already visited
  • Output: (B, 98) logits, row-major camera index row * 14 + col
  • Validation at the released epoch (47): top-1 0.830, top-3 0.957, P@10 0.904
  • Original run name: reward_UA_sigz_final_v3
python 4_policy/train_reward_UA.py \
    --reward_roots /result/APOBU/reward_dataset/ua_reward_dataset_ycb_v3_unified \
    --save_dir /result/DiVE/reward_UA_sigz_final_v3 \
    --wandb_run_name reward_UA_sigz_final_v3 \
    --loss_type bce_sigmoid_z \
    --temperature 1.0 \
    --use_seen_mask \
    --device 1

push_selection_policy β€” push selection policy

  • Model: 4_policy_test/model_reward_UB.py β†’ RewardNetUB (self-contained, no support_code dependency; not the 29-channel variant in support_code/model_reward_UB.py)
  • Trainer: 4_policy_test/train_reward_UB.py
  • Architecture: siamese 3D CNN over concat(belief, swept_map_i) with an overlap-residual score and a transformer block attending across candidate actions (base_ch=16, beta=0.5, use_attn=True, 4 heads), 1.43M parameters
  • Input: belief map (B, 1, 60, 120, 80) and candidate swept maps (B, N, 60, 120, 80)
  • Output: (B, N) scores over the candidate pushes
  • Validation at the released epoch (21): top-1 0.336, top-3 0.643, P@10 0.853
  • This is the exact checkpoint used for the evaluation runs reported in the paper. A later epoch-28 checkpoint of the same run exists (marginally lower val loss, top-1 0.356 / top-3 0.629); the epoch-21 weights are released so the reported numbers reproduce
  • Original run name: reward_UB_full_siamese_sigz_b05_attn_final_v3
python 4_policy_test/train_reward_UB.py \
    --reward_roots /result/APOBU/reward_dataset/ub_reward_dataset_ycb_v3_unified \
    --data_roots /data/APOBU/beliefmap_high_occlusion_ycb_v3 \
                 /result/DiVE_data/beliefmap_low_occlusion_ycb_v3 \
    --save_dir /result/APOBU/DiVE/reward_UB_full_siamese_sigz_b05_attn_final_v3 \
    --wandb_run_name reward_UB_full_siamese_sigz_b05_attn_final_v3 \
    --loss_type bce_sigmoid_z \
    --temperature 1.0 \
    --beta 0.5 \
    --use_attn \
    --attn_heads 4 \
    --max_steps_per_epoch 3000 \
    --max_val_steps 500 \
    --batch_size 8 \
    --epochs 50 \
    --patience 6 \
    --device 2,3

belief_map_estimator β€” recursive semantic belief updater

  • Model: BeliefUpdateNetRecursive(n_sem=15, base_ch=32) over a 3D U-Net backbone, 14.70M parameters
  • Input: 34 channels = [alpha/50, beta/50, occ, swept_map, one_hot(sem_obs) x15, sem_belief x15]
  • Output: 17 channels = [new_alpha, new_beta, sem_logits x15]; the softmaxed semantic logits and the Beta parameters are fed back at the next step (recursive)
  • YCB v3 uses classes 0-11 (12 active); slots 12-14 are reserved, hence n_sem=12 at evaluation with n_sem_model=15
  • Stored as a plain state_dict (58.8 MB):
import torch
from estimator import BeliefUpdateNetRecursive
net = BeliefUpdateNetRecursive(n_sem=15, base_ch=32)
net.load_state_dict(torch.load("belief_map_estimator/estimator.pt", map_location="cpu"))
  • Trained from scratch with DDP on 4 GPUs (--batch 32, --n_sem 15) on beliefmap_{high,low}_occlusion_ycb_v3
  • Reference evaluation, DiVE policy with this estimator as the central tracker, low-occlusion test set, 100 episodes x 25-step budget: mIoU 0.9279, occupancy IoU 0.8721, semantic mIoU 0.8116, 4.92 pushes per episode

Training data

Cluttered shelf scenes generated in NVIDIA Isaac Sim with YCB objects, at low and high occlusion. Extreme-occlusion scenes are held out for evaluation only. The dataset release is in preparation; see the project page.

Citation

@inproceedings{lee2026dive,
  title     = {Learning Decomposed Visibility for Efficient Active Exploration of Cluttered Scenes},
  author    = {Lee, Suyun and Choi, Minsoo and Gong, Jihwan and Nam, Unghui and Bae, Minji and Shim, Byonghyo},
  booktitle = {Conference on Robot Learning (CoRL)},
  year      = {2026},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading