Spatial-Interactor-Qwen3-VL-8B

This is a full-parameter BF16 checkpoint from Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World, based on Qwen/Qwen3-VL-8B-Instruct.

Spatial-Interactor learns local world-state and ego-motion transitions through supervised fine-tuning, then applies On-Policy Distillation (OPD) to integrate successive transitions over long trajectories. The privileged transition trace is used only during training. At inference, this checkpoint takes the same image/video and question inputs as its base model and requires no extra trace, reward model, or teacher branch.

Resources

Usage

Use the standard Transformers inference interface documented for the base model. Load this repository in place of the base model identifier:

model_id = "kagakouko/Spatial-Interactor-Qwen3-VL-8B"

For video evaluation, preserve chronological frame order and use the frame budget specified by the target benchmark. The paper's main video evaluation uses 32 ordered frames.

Training summary

The SFT stage trains on the reported L1-L2 split of LSI-108K together with the public spatial QA mixture described in the paper. OPD starts from that SFT checkpoint and combines verifiable answer rewards with CoT-only privileged self-distillation on long-horizon video questions. The visual encoder remains frozen while the language model and multimodal projector are updated.

License

This checkpoint is released under Apache-2.0, following the base model license. Users must separately comply with licenses and terms governing input datasets and media.

Downloads last month
12
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kagakouko/Spatial-Interactor-Qwen3-VL-8B

Finetuned
(585)
this model
Quantizations
1 model

Collection including kagakouko/Spatial-Interactor-Qwen3-VL-8B