User-Feedback-Driven Adaptation for Vision-and-Language Navigation
Data and checkpoint for the paper User-Feedback-Driven Adaptation for Vision-and-Language Navigation, IEEE Transactions on Multimedia, 2026.
[Paper (IEEE TMM)] [arXiv] [Code]
Download
huggingface-cli download Peachilk/UFD --local-dir data
This gives the data/ directory used by the scripts in the code repository. Some files come from other projects and are not included here: the connectivity graphs and CLIP features (ScaleVLN), prevalent_aug_train_enc.json (GSA-R2R) and the initial GR-DUET checkpoint best_val_unseen (GSA-VLN). The code README explains where to get them.
Contents
| Path | Description |
|---|---|
ckpts/best_overall_iter_23500 |
GR-DUET fine-tuned on paths reconstructed from user feedback in the Basic setting (rec_basic); the checkpoint with the best average validation SPL (iteration 23,500). It is also the model evaluated with the memory bank. |
split_basic/split_test_500_<scan>.json |
500 instructions per scan on which user feedback is collected; input to path reconstruction (24 scans, 12,000 entries). |
split_mixed/split_test_500_<scan>.json |
Same for the mixed user styles (21 scans, 10,500 entries). |
rec_basic/merged_train_dataset.json |
Training set of reconstructed paths with 5 to 7 viewpoints (24 scans, 11,342 entries). |
rec_mixed/merged_train_dataset.json |
Same for the mixed user styles (21 scans, 9,766 entries). |
rec_basic/merged_500_test_dataset.json, rec_mixed/merge_500_test_dataset_mixed.json |
The 500 instructions of all scans in one file, used to build the memory bank. |
GSA_Dataset/Validation/Residential/Reconstruction/validation_set_100.json |
Validation set: up to 100 reconstructed paths per scan (24 scans, 2,390 entries). |
GSA_Dataset/Test/Residential/Reconstruction/merged_test_dataset.json |
Test set: 100 held-out instructions per scan (24 scans, 2,400 entries). |
GSA_Dataset/Test/Residential/Recmixed/merged_test_dataset.json |
Test set for the mixed user styles (21 scans, 2,100 entries). |
GSA_Dataset/Train/R2R_train_enc.json |
R2R training instructions with BERT token ids. |
scanvp_cands_relangles_with_habitat.json |
Candidate viewpoints and relative angles of every viewpoint. |
pano_inputs_habitats_more_scan.h5 |
Precomputed panorama inputs of every viewpoint, used to build full graphs during training. |
basic uses the Basic instructions of GSA-R2R; mixed combines instructions in five user styles (child, keith, moira, rachel, sheldon).
License
The data and checkpoint are for non-commercial research use only. Please also follow the terms of use of the original datasets: R2R, Matterport3D, HM3D and GSA-R2R. The code is released under the MIT License.
Citation
@article{yu2026userfeedback,
title = {User-Feedback-Driven Adaptation for Vision-and-Language Navigation},
author = {Yu, Yongqiang and Li, Xuhui and Mahmood, Hazza and Zhou, Jinxing and Hong, Haodong and Jiang, Longtao and Xu, Zhiqiang and Wu, Qi and Chang, Xiaojun},
journal = {IEEE Transactions on Multimedia},
year = {2026},
doi = {10.1109/TMM.2026.3724731}
}