Qwen3-VL-8B-TGRPO: TGRPO baseline

This repository contains the Qwen3-VL-8B weights fine-tuned with TGRPO, released as a baseline for the paper Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning.

Overview

This checkpoint is fine-tuned on top of Qwen3-VL with TGRPO (Temporal GRPO), one of the RL baselines compared in the CRPO paper.

Video large language models (Video LLMs) often answer video questions through shortcuts such as single-frame cues and language priors rather than by tracking spatiotemporal dynamics. The CRPO paper studies this problem and compares against multiple RL baselines, of which this checkpoint is one.

Resources

Evaluation

The model is evaluated on DyBench, a paired counterfactual video benchmark with 3,014 videos covering reversible dynamics, moving direction, and event sequence, together with standard video-QA benchmarks (Video-MME, TempCompass, MVBench, TimeBlind). See the paper for full numbers and comparison with CRPO.

Downloads last month
15
Safetensors
Model size
770k params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for ddz16/Qwen3-VL-8B-TGRPO