YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Enhanced T-GRPO Phase 1

Project: CS41 Enhanced T-GRPO for Video Temporal Reasoning (USYD Capstone)

Model Details

  • Base model: Qwen2.5-VL-7B-COT-SFT
  • Algorithm: Enhanced T-GRPO with per-sample margin reward, 6 corruption types, curriculum learning
  • Training data: 500 video samples (CLEVRER, STAR, PerceptionTest)
  • Training steps: 200 (50 steps Run 1 + 150 steps from checkpoint-50)
  • GPU: 2ร—A100 SXM 80GB (ZeRO-3 + CPU offload)

Algorithm Modifications (vs Original T-GRPO)

  1. Per-sample continuous margin reward replacing binary batch-level reward
  2. 6 corruption types (shuffle, reverse, mask, chunk_swap, speed, loop) replacing single shuffle
  3. Curriculum learning โ€” corruption strength ramps from 0.3 to 1.0
  4. Equalized generation count โ€” same G for both ordered and shuffled conditions

Training Config

Parameter Value
num_generations 4
max_pixels 200704
max_prompt_length 8192
learning_rate 1e-6
beta (KL coeff) 0.04
corruption_strength 1.0 (with curriculum)
margin_scale 0.5
reward_threshold 0.1
gradient_accumulation 2
DeepSpeed ZeRO-3 + CPU offload

Training Curves

  • Reward: 1.5 โ†’ 1.8-2.0 (U-shaped, consistent with curriculum learning)
  • Temporal rewards: Non-zero in 88% of steps (mean 0.077)
  • KL divergence: Healthy range, peaked at 0.011
  • All 6 corruption types triggered successfully

WandB: https://wandb.ai/wuguangbo464-the-university-of-sydney/huggingface/runs/lep2va20

Evaluation

Benchmark Score
MMVU (mc) 59.0%

Comparison

Model MMVU (mc)
Qwen2.5-VL-7B-SFT (starting point) 61.3%
Baseline T-GRPO (P0 fixes only) 61.1%
Enhanced T-GRPO (this model) 59.0%
Video-R1-7B (paper, full data) 64.2%

Notes

Phase 1 small-scale validation. MMVU score drop is expected โ€” trained on 0.2% of full data, and MMVU tests domain knowledge, not temporal reasoning. Training curves confirm the algorithm works (reward trending up, temporal signal active). Phase 2 will use full 260k dataset and evaluate on VSI-Bench/TempCompass (temporal reasoning benchmarks).

Downloads last month
9
Safetensors
Model size
8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support