YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Enhanced T-GRPO Phase 1
Project: CS41 Enhanced T-GRPO for Video Temporal Reasoning (USYD Capstone)
Model Details
- Base model: Qwen2.5-VL-7B-COT-SFT
- Algorithm: Enhanced T-GRPO with per-sample margin reward, 6 corruption types, curriculum learning
- Training data: 500 video samples (CLEVRER, STAR, PerceptionTest)
- Training steps: 200 (50 steps Run 1 + 150 steps from checkpoint-50)
- GPU: 2รA100 SXM 80GB (ZeRO-3 + CPU offload)
Algorithm Modifications (vs Original T-GRPO)
- Per-sample continuous margin reward replacing binary batch-level reward
- 6 corruption types (shuffle, reverse, mask, chunk_swap, speed, loop) replacing single shuffle
- Curriculum learning โ corruption strength ramps from 0.3 to 1.0
- Equalized generation count โ same G for both ordered and shuffled conditions
Training Config
| Parameter | Value |
|---|---|
| num_generations | 4 |
| max_pixels | 200704 |
| max_prompt_length | 8192 |
| learning_rate | 1e-6 |
| beta (KL coeff) | 0.04 |
| corruption_strength | 1.0 (with curriculum) |
| margin_scale | 0.5 |
| reward_threshold | 0.1 |
| gradient_accumulation | 2 |
| DeepSpeed | ZeRO-3 + CPU offload |
Training Curves
- Reward: 1.5 โ 1.8-2.0 (U-shaped, consistent with curriculum learning)
- Temporal rewards: Non-zero in 88% of steps (mean 0.077)
- KL divergence: Healthy range, peaked at 0.011
- All 6 corruption types triggered successfully
WandB: https://wandb.ai/wuguangbo464-the-university-of-sydney/huggingface/runs/lep2va20
Evaluation
| Benchmark | Score |
|---|---|
| MMVU (mc) | 59.0% |
Comparison
| Model | MMVU (mc) |
|---|---|
| Qwen2.5-VL-7B-SFT (starting point) | 61.3% |
| Baseline T-GRPO (P0 fixes only) | 61.1% |
| Enhanced T-GRPO (this model) | 59.0% |
| Video-R1-7B (paper, full data) | 64.2% |
Notes
Phase 1 small-scale validation. MMVU score drop is expected โ trained on 0.2% of full data, and MMVU tests domain knowledge, not temporal reasoning. Training curves confirm the algorithm works (reward trending up, temporal signal active). Phase 2 will use full 260k dataset and evaluate on VSI-Bench/TempCompass (temporal reasoning benchmarks).
- Downloads last month
- 9
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support