NavGPT3-8B
NavGPT3-8B is the 8B NavGPT VLA, the low-level navigation policy of NavGPT-3. Given a natural-language instruction and a history of four-view RGB observations (front, right, back, left), it predicts the next eight waypoints. In NavGPT-3, a Planner chooses the instructions and decides when to stop; the VLA carries them out.
Resources
- Paper: NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime
- Code: metacognitionai/NavGPT-3
- Other size: NavGPT3-4B
Model details
| Base model | Qwen3-VL-8B-Instruct |
| Parameters | 8.77B, BF16 |
| Action head | Two-layer MLP on the final hidden state of the last prompt token; outputs 8 waypoints of (x, y, θ), normalised |
| Observations | Up to 16 history steps × 4 views, 400×400 at 120° field of view |
| Visual tokens | One budget of 3,072 tokens shared by all images, split per image by recency, camera and how much the frame changed |
| Training data | About 19.28M examples: vision-and-language navigation, embodied navigation QA, cross-embodiment navigation and grounding |
| Code | NavGPT3 classes on Transformers 5.18 |
Results
Vision-and-language navigation on the R2R-CE and RxR-CE Val-Unseen splits, as reported in the NavGPT-3 paper. The NavGPT3 rows are the model on its own, without the NavGPT-3 Planner.
| Method | Size | R2R NE↓ | R2R OSR↑ | R2R SR↑ | R2R SPL↑ | RxR NE↓ | RxR nDTW↑ | RxR SR↑ | RxR SPL↑ |
|---|---|---|---|---|---|---|---|---|---|
| NavFoM | 7B | 4.61 | 72.1 | 61.7 | 55.3 | 4.74 | 65.8 | 64.4 | 56.2 |
| ABot-N0 | 4B | 3.78 | 70.8 | 66.4 | 63.9 | 3.83 | – | 69.3 | 60.0 |
| AstraNav-World | 3B+5B | 3.86 | 73.9 | 67.9 | 65.4 | 3.82 | – | 72.9 | 61.5 |
| OmniNav | 3B | 3.74 | 74.6 | 69.5 | 66.1 | 3.77 | – | 73.6 | 62.0 |
| Qwen-RobotNav | 4B | 3.80 | 77.2 | 69.5 | 63.6 | 3.80 | 71.9 | 75.2 | 65.0 |
| Qwen-RobotNav | 8B | 3.53 | 78.5 | 72.1 | 66.6 | 3.58 | 72.5 | 76.5 | 65.7 |
| Robostral Navigate | 8B | 3.20 | 81.3 | 77.4 | 74.2 | 3.47 | – | 75.1 | 68.7 |
| NavGPT3-4B | 4B | 3.41 | 79.91 | 72.54 | 67.19 | 3.35 | 73.33 | 76.77 | 67.77 |
| NavGPT3-8B | 8B | 3.29 | 80.19 | 74.51 | 68.54 | 3.05 | 74.85 | 78.19 | 68.98 |
NE is navigation error in metres; OSR (oracle success), SR (success), SPL (success weighted by path length) and nDTW (path fidelity) are percentages.
In the complete NavGPT-3 system, where a Planner directs NavGPT3-8B, success rises to 81.51 SR on R2R-CE and 90.43 SR on RxR-CE.
Usage
The model needs the NavGPT3 classes from the
NavGPT-3 repository; it does not
load with AutoModel alone.
hf download Metacognition-AI/NavGPT3-8B --local-dir checkpoints/NavGPT3-8B
from navgpt.vla.runtime import load_from_pretrained
model, processor = load_from_pretrained("checkpoints/NavGPT3-8B", device="cuda")
model.predict_actions(**inputs) returns the normalised waypoints for processor
outputs. The navgpt.vla agent keeps the multi-view history, assigns the
per-image token budgets and scales the waypoints to metres and radians. To serve
the model for Habitat evaluation, set its directory in configs/paths.yaml and run
python -m navgpt.vla --config configs/experiments/vln/vla_8b_r2r.yaml.
Limitations
- Trained and evaluated in Habitat simulation (R2R-CE, RxR-CE). Waypoint scales follow the simulated embodiment; other robots need their own scales and testing.
- Evaluated on English instructions.
License
The weights are released under the GNU Affero General Public License v3.0. They are fine-tuned from Qwen3-VL-8B-Instruct, released by the Qwen team under the Apache License 2.0; see NOTICE. The NavGPT-3 source code is licensed separately.
Citation
@article{zhou2026navgpt3,
title={NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime},
author={Zhou, Gengze and Hong, Yicong and Zhang, Jiazhao and Zhao, Xunyi and Zhou, Jian and Lei, Zixing and Wang, Zun and Zhao, Chongyang and Chen, Xionghui and Gould, Stephen and van den Hengel, Anton and Wu, Qi},
journal={arXiv preprint arXiv:2610.10787},
year={2026}
}
- Downloads last month
- 12
Model tree for Metacognition-AI/NavGPT3-8B
Base model
Qwen/Qwen3-VL-8B-Instruct