Image-Text-to-Text
Transformers
Safetensors
qwen3_vl
conversational

SpatialBlock-4B-reason

This repository contains the SpatialBlock-4B-reason checkpoint from the paper SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem.

It is a fine-tuned version of Qwen3-VL-4B-Instruct on the synthetic SpatialBlock-15k dataset. The model directly predicts answers to spatial reasoning tasks such as 3D-to-2D projection, viewpoint transformation, and structural combination.

For training details, evaluation results, and the companion “direct” model, please refer to the GitHub repository: https://github.com/rsoohyun/SpatialBlock.

Downloads last month
23
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for rsoohyun/SpatialBlock-4B-reason

Finetuned
(429)
this model

Dataset used to train rsoohyun/SpatialBlock-4B-reason

Collection including rsoohyun/SpatialBlock-4B-reason

Paper for rsoohyun/SpatialBlock-4B-reason