ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
Abstract
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely semantic-centric, and often fail to reliably aggregate consistent spatial evidence from redundant video observations, leading to inefficient or unstable reasoning. To address these issues, we propose ConsiSpace, a geometry-consistency-aware framework for geometry-sensitive video spatial reasoning that turns spatial consistency into both an evidence organization principle and an explicit post-SFT learning signal. We build a geometry-consistent memory (GCM) including implicit evidence tokens and explicit geometric cues, and leverage efficient organization strategies to compactly preserve task-related spatial evidence. Furthermore, we utilize unified consistency self-supervised reinforcement learning (UC-SSRL) after supervised fine-tuning to improve cross-view stability, with answer-, metric-, and topology-consistency rewards. Extensive experiments on three spatial-reasoning benchmarks, VSI-Bench, OSI-Bench, and MMSI-Video-Bench, show consistent gains, improving the average score by 12.6 points over the strongest baselines.
Community
We propose a geometry-consistency-aware framework for video spatial reasoning.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning (2026)
- LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video (2026)
- Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views (2026)
- OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping (2026)
- Dual-Pathway Geometry-Aware MLLM for Spatial Intelligence (2026)
- Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models (2026)
- Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.17599 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper