Spatial-Interactor visual presentation cover

Spatial-Interactor Qwen2.5-VL-7B

From local state transitions to long-horizon spatial reasoning

Project page Paper PDF Code Dataset

This is the full-parameter BF16 Spatial-Interactor checkpoint based on Qwen/Qwen2.5-VL-7B-Instruct. It learns local world-state and ego-motion transitions through supervised fine-tuning, then uses On-Policy Distillation (OPD) to integrate successive transitions over long trajectories.

The privileged transition trace is used only during training. At inference, this checkpoint takes the same image/video and question inputs as its base model, with no extra trace, reward model, or teacher branch.

How Spatial-Interactor learns

On-Policy Distillation pipeline

L1 and L2 establish local state-transition modeling. On L3, verifiable answer rewards supervise the result while same-prefix privileged distillation guides the intermediate reasoning process.

Usage

Use the standard Transformers interface for the base model and load this repository in place of the base identifier:

model_id = "kagakouko/Spatial-Interactor-Qwen2.5-VL-7B"

For video evaluation, preserve chronological frame order and use the frame budget specified by the target benchmark. The paper's main video evaluation uses 32 ordered frames.

Training summary

The SFT stage uses the reported L1-L2 split of LSI-108K together with the public spatial QA mixture described in the paper. OPD starts from that SFT checkpoint and combines verifiable answer rewards with CoT-only privileged self-distillation on long-horizon video questions. The visual encoder remains frozen while the language model and multimodal projector are updated.

License

This checkpoint is released under Apache-2.0, following the base model license. Users must also comply with licenses and terms governing input datasets and media.

Downloads last month
20
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kagakouko/Spatial-Interactor-Qwen2.5-VL-7B

Finetuned
(1222)
this model
Quantizations
1 model

Collection including kagakouko/Spatial-Interactor-Qwen2.5-VL-7B