GRASP-Qwen3-VL-8B
📜 Paper · 🌐 Project Page · 💻 Code · 📦 Dataset
Qwen3-VL-8B post-trained on GRASP with Social Grounding Reward (SGR) — a GRPO learning signal that rewards reasoning about the correct participants in each gaze/gesture interaction, built on GRASP's structured social events.
Introduced in GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions (NeurIPS 2026).
Usage
The model answers social reasoning questions over multi-person video, producing <gaze>/<gesture> grounding tags and a final <answer> tag.
from transformers import AutoModelForImageTextToText, AutoProcessor
model = AutoModelForImageTextToText.from_pretrained(
"interlive/GRASP-Qwen3-VL-8B", torch_dtype="bfloat16", device_map="auto"
)
processor = AutoProcessor.from_pretrained("interlive/GRASP-Qwen3-VL-8B")
Evaluation on GRASP-Bench: see eval/eval_social_qa.py in the code release. Videos are sampled at 2 fps.
License
The model weights are released under CC BY-NC 4.0 for non-commercial research use only. The base model, Qwen3-VL-8B-Instruct, is licensed under Apache 2.0.
Citation
@article{kim2026grasp,
title={GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions},
author={Kim, Junho and Cao, Xu and Yang, Houze and Boote, Bikram and Jojic, Ana and Ryan, Fiona and Lai, Bolin and Lee, Sangmin and Rehg, James M},
journal={arXiv preprint arXiv:2605.15764},
year={2026}
}
- Downloads last month
- 20
Model tree for interlive/GRASP-Qwen3-VL-8B
Base model
Qwen/Qwen3-VL-8B-Instruct