GRASP-Qwen3-VL-8B

📜 Paper · 🌐 Project Page · 💻 Code · 📦 Dataset

Qwen3-VL-8B post-trained on GRASP with Social Grounding Reward (SGR) — a GRPO learning signal that rewards reasoning about the correct participants in each gaze/gesture interaction, built on GRASP's structured social events.

Introduced in GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions (NeurIPS 2026).

Usage

The model answers social reasoning questions over multi-person video, producing <gaze>/<gesture> grounding tags and a final <answer> tag.

from transformers import AutoModelForImageTextToText, AutoProcessor

model = AutoModelForImageTextToText.from_pretrained(
    "interlive/GRASP-Qwen3-VL-8B", torch_dtype="bfloat16", device_map="auto"
)
processor = AutoProcessor.from_pretrained("interlive/GRASP-Qwen3-VL-8B")

Evaluation on GRASP-Bench: see eval/eval_social_qa.py in the code release. Videos are sampled at 2 fps.

License

The model weights are released under CC BY-NC 4.0 for non-commercial research use only. The base model, Qwen3-VL-8B-Instruct, is licensed under Apache 2.0.

Citation

@article{kim2026grasp,
  title={GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions},
  author={Kim, Junho and Cao, Xu and Yang, Houze and Boote, Bikram and Jojic, Ana and Ryan, Fiona and Lai, Bolin and Lee, Sangmin and Rehg, James M},
  journal={arXiv preprint arXiv:2605.15764},
  year={2026}
}
Downloads last month
20
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for interlive/GRASP-Qwen3-VL-8B

Finetuned
(623)
this model

Paper for interlive/GRASP-Qwen3-VL-8B