Checkpoints β Projection-Assisted Segmentation Reproduction
Model weights for reproducing projection-assisted gaze-prompted segmentation on the Aria Digital Twin dataset.
Companion data: tianxia2/projseg-adt-seq144-subset.
Files
| File | Size | What |
|---|---|---|
efficient_sam_vitt.pt |
40 MB | EfficientSAM ViT-Tiny. The anchor segmenter, prompted with a single gaze point. Unmodified upstream weights from yformer/EfficientSAM (Apache-2.0), redistributed for convenience. |
refinenetgaze.pth |
5.9 MB | RefineNetGaze. Optional U-Net that refines a reused mask from RGB + prior mask + gaze heatmap (5-channel input, base=16). Trained by us on ADT. |
Usage
huggingface-cli download tianxia2/projseg-checkpoints --local-dir checkpoints
The reproduction artifact expects them at checkpoints/ (override with
PROJSEG_CKPT_ROOT).
Notes
refinenetgaze.pthis stored as a plain tensor state dict, so it loads under PyTorch β₯ 2.6's defaultweights_only=True. Construct the model asRefineNetGaze(base=16).- Refinement is optional and off by default. On the reported EfficientSAM +
dense_gtconfiguration it lowers mIoU (0.3283 β 0.3111), because it was trained against a different anchor-mask distribution than the one ESAM anchors produce. It is shipped for completeness and ablation, not as part of the headline result. - EfficientSAM is class-agnostic. Prompted with one gaze point it segments some coherent region around that point, which does not always coincide with the annotated ADT instance β hence modest absolute mIoU on this benchmark.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support