NavGPT3-8B

NavGPT3-8B is the 8B NavGPT VLA, the low-level navigation policy of NavGPT-3. Given a natural-language instruction and a history of four-view RGB observations (front, right, back, left), it predicts the next eight waypoints. In NavGPT-3, a Planner chooses the instructions and decides when to stop; the VLA carries them out.

Resources

Model details

Base model Qwen3-VL-8B-Instruct
Parameters 8.77B, BF16
Action head Two-layer MLP on the final hidden state of the last prompt token; outputs 8 waypoints of (x, y, θ), normalised
Observations Up to 16 history steps × 4 views, 400×400 at 120° field of view
Visual tokens One budget of 3,072 tokens shared by all images, split per image by recency, camera and how much the frame changed
Training data About 19.28M examples: vision-and-language navigation, embodied navigation QA, cross-embodiment navigation and grounding
Code NavGPT3 classes on Transformers 5.18

Results

Vision-and-language navigation on the R2R-CE and RxR-CE Val-Unseen splits, as reported in the NavGPT-3 paper. The NavGPT3 rows are the model on its own, without the NavGPT-3 Planner.

Method Size R2R NE↓ R2R OSR↑ R2R SR↑ R2R SPL↑ RxR NE↓ RxR nDTW↑ RxR SR↑ RxR SPL↑
NavFoM 7B 4.61 72.1 61.7 55.3 4.74 65.8 64.4 56.2
ABot-N0 4B 3.78 70.8 66.4 63.9 3.83 – 69.3 60.0
AstraNav-World 3B+5B 3.86 73.9 67.9 65.4 3.82 – 72.9 61.5
OmniNav 3B 3.74 74.6 69.5 66.1 3.77 – 73.6 62.0
Qwen-RobotNav 4B 3.80 77.2 69.5 63.6 3.80 71.9 75.2 65.0
Qwen-RobotNav 8B 3.53 78.5 72.1 66.6 3.58 72.5 76.5 65.7
Robostral Navigate 8B 3.20 81.3 77.4 74.2 3.47 – 75.1 68.7
NavGPT3-4B 4B 3.41 79.91 72.54 67.19 3.35 73.33 76.77 67.77
NavGPT3-8B 8B 3.29 80.19 74.51 68.54 3.05 74.85 78.19 68.98

NE is navigation error in metres; OSR (oracle success), SR (success), SPL (success weighted by path length) and nDTW (path fidelity) are percentages.

In the complete NavGPT-3 system, where a Planner directs NavGPT3-8B, success rises to 81.51 SR on R2R-CE and 90.43 SR on RxR-CE.

Usage

The model needs the NavGPT3 classes from the NavGPT-3 repository; it does not load with AutoModel alone.

hf download Metacognition-AI/NavGPT3-8B --local-dir checkpoints/NavGPT3-8B
from navgpt.vla.runtime import load_from_pretrained

model, processor = load_from_pretrained("checkpoints/NavGPT3-8B", device="cuda")

model.predict_actions(**inputs) returns the normalised waypoints for processor outputs. The navgpt.vla agent keeps the multi-view history, assigns the per-image token budgets and scales the waypoints to metres and radians. To serve the model for Habitat evaluation, set its directory in configs/paths.yaml and run python -m navgpt.vla --config configs/experiments/vln/vla_8b_r2r.yaml.

Limitations

  • Trained and evaluated in Habitat simulation (R2R-CE, RxR-CE). Waypoint scales follow the simulated embodiment; other robots need their own scales and testing.
  • Evaluated on English instructions.

License

The weights are released under the GNU Affero General Public License v3.0. They are fine-tuned from Qwen3-VL-8B-Instruct, released by the Qwen team under the Apache License 2.0; see NOTICE. The NavGPT-3 source code is licensed separately.

Citation

@article{zhou2026navgpt3,
  title={NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime},
  author={Zhou, Gengze and Hong, Yicong and Zhang, Jiazhao and Zhao, Xunyi and Zhou, Jian and Lei, Zixing and Wang, Zun and Zhao, Chongyang and Chen, Xionghui and Gould, Stephen and van den Hengel, Anton and Wu, Qi},
  journal={arXiv preprint arXiv:2610.10787},
  year={2026}
}
Downloads last month
12
Safetensors
Model size
9B params
Tensor type
BF16
·
Video Preview
loading

Model tree for Metacognition-AI/NavGPT3-8B

Finetuned
(623)
this model

Paper for Metacognition-AI/NavGPT3-8B