YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
The Earth in One Gaze: Training-Free
Active Focus for UHR Remote Sensing Understanding
Yao Zhang1,*, Pengyu Dai2,3,*, Wei Guo1,β , Jian Liang1,
Jian Song3, Yafei Ou3, Hongruixuan Chen3,β , Naoto Yokoya2,3
1 Wuhan University 2 The University of Tokyo 3 RIKEN AIP
* Equal contribution β Corresponding authors
Abstract
Multimodal large language models must balance local detail against scene context when interpreting ultra-high-resolution remote-sensing imagery within a limited visual-input budget. Knowing where to look is not enough: what a model can infer from selected evidence also depends on how that evidence is presented. We introduce GazeEarth, a training-free framework that couples question-guided region selection with full-scene foveated observation. A frozen model selects evidence cells from an indexed overview. A deterministic, topology-preserving warp then resamples the original image onto a fixed-size canvas, enlarging the selected neighborhood while compressing the periphery. The same model answers from this focused view, retaining local evidence within its surrounding scene context. GazeEarth requires at most two model calls, with no fine-tuning, external selector, or iterative search.
Environment
Python 3.11 and a CUDA GPU are recommended. Run the following commands from the repository root:
conda create -n gazeearth python=3.11 -y
conda activate gazeearth
pip install torch==2.6.0 torchvision==0.21.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt
pip install -e . --no-deps
python -m nltk.downloader wordnet omw-1.4
Data
Arrange LRS-GRO, XLRS-Bench, and MME-RealWorld-RS as follows:
data/
βββ LRS-GRO/
β βββ test.json
β βββ images/
β βββ ...
βββ XLRS-Bench/
β βββ test.jsonl
β βββ images/
β βββ ...
βββ MME-RealWorld-RS/
βββ test.jsonl
βββ images/
βββ ...
For custom paths, update dataset.annotation_path and dataset.root in the YAML config.
Backbone Configuration
Use the default Qwen3-VL configuration, or add a model configuration to either inference or evaluation:
| Backbone | Model configuration |
|---|---|
| Qwen3-VL-8B-Instruct | Default in the benchmark configs |
| LLaVA-v1.6-Mistral-7B | --model-config configs/llava.yaml |
| Intern-S1-mini | --model-config configs/intern_s1.yaml |
| GPT-4o via OpenRouter | --model-config configs/gpt4o.yaml |
For GPT-4o, set the OPENROUTER_API_KEY environment variable before running. API usage is billed by the provider.
Evaluation
# Evaluate a benchmark.
python scripts/eval.py --config configs/xlrs_bench.yaml --output outputs/xlrs_qwen
python scripts/eval.py --config configs/lrs_gro.yaml
python scripts/eval.py --config configs/mme_realworld_rs.yaml
# For the same command and output directory:
# append --resume to continue an interrupted run;
# append --overwrite to back up existing outputs and restart.
# Ask a question about your own image.
python scripts/infer.py --images path/to/image.jpg --question "What is shown?"
Outputs include predictions.jsonl, traces.jsonl, results.json, and run_manifest.json. Do not combine --resume and --overwrite.
Citation
If you find GazeEarth useful in your research, please consider citing our paper:
@misc{zhang2026gazeearth,
title = {The Earth in One Gaze: Training-Free Active Focus for {UHR} Remote Sensing Understanding},
author = {Yao Zhang and Pengyu Dai and Wei Guo and Jian Liang and Jian Song and Yafei Ou and Hongruixuan Chen and Naoto Yokoya},
year = {2026},
eprint = {2609.31747},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.31747}
}
Acknowledgements
We thank the teams behind Qwen3-VL, LLaVA, Intern-S1, and GPT-4o for making their models accessible, and the creators of LRS-GRO, XLRS-Bench, and MME-RealWorld for providing the benchmarks used in this work.
