SPARC-Qwen3.5-0.8B-VTFT

Qwen3.5-0.8B fully fine-tuned for embodied spatial reasoning using VQA data generated from SPARC annotations. VTFT denotes vision-tower fine-tuning.

Training data

The training mixture contains SPARC-generated VQA data from ours_adaptive_det_soft_snr_sp8, FSD, RoboPoint, and LLaVA-OneVision2. SPARC samples use an annotation-quality threshold of 0.97, are sorted by score, and are capped at 700 samples per object. This is the same mixture as the 4B release.

Release SPARC VQA (filtered) FSD RoboPoint LLaVA-OneVision2 EO-1.5M
Qwen3.5-4B Yes Yes Yes Yes No
Qwen3.5-0.8B-VTFT Yes Yes Yes Yes No
Qwen3.5-9B-EO Yes Yes Yes Yes Yes

The SPARC training data is the filtered release subset, which needs no further SPARC filtering. Its filtering script and release mixture manifest are in irl-kit/SPARC-VQA-Raw. FSD, RoboPoint, LLaVA-OneVision2, and EO-1.5M remain their respective upstream datasets.

Prompting

Prompt formatting is important for these models. Use the bundled chat_template.jinja through processor.apply_chat_template(..., add_generation_prompt=True), with one user turn containing the image(s) followed by the text question. Disable thinking/reasoning mode to match evaluation.

For a single point, append exactly:

Output the point coordinates in JSON format like [{"point_2d": [x, y], "label": "target"}]. Use integer coordinates between 0 and 1000.

For a trajectory or multiple points, append exactly:

Return only a JSON list like [{"point_2d": [x1, y1], "label": "point_1"}, {"point_2d": [x2, y2], "label": "point_2"}, ...]. Use integer coordinates between 0 and 1000.

Training

Both the vision encoder and vision projector are trainable. The model was fully fine-tuned for one epoch with a learning rate of 2e-5 and a maximum sequence length of 5600.

Evaluation

This is the strongest evaluated 0.8B variant: local full-benchmark aggregate 0.605, compared with 0.596 for the frozen-vision 0.8B run. The paper appendix reports a 53.2 pointing/VQA average for this variant.

Model Aggregate Where2Place RefSpatial location IA-Bench RoboRefIt testA VA Bench-P
Qwen3.5-4B 0.698 72.0 59.0 79.0 85.7 65.7
Qwen3.5-0.8B-VTFT 0.605 58.0 47.0 76.7 80.9 48.3
Qwen3.5-9B-EO 0.719 76.0 68.0 78.5 85.2 68.7

Citation

@article{blank2026sparc,
  title={SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale},
  author={Blank, Nils and others},
  journal={arXiv preprint arXiv:2606.13497},
  year={2026}
}
Downloads last month
145
Safetensors
Model size
0.9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for irl-kit/SPARC-Qwen3.5-0.8B-VTFT