LLaVA-Video-72B SIMS-V 3Q-25K

This is a full-parameter fine-tune of lmms-lab/LLaVA-Video-72B-Qwen2 on the 25K-example SIMS-V 3Q spatial-reasoning mixture.

Lineage

The training mixture contains 25,000 programmatically generated examples:

Task Examples
Absolute distance, open-ended 8,616
Appearance order, multiple-choice 6,384
Relative direction, medium 3,905
Relative direction, hard 3,853
Relative direction, easy 2,242

Training recipe

  • Full-parameter tuning of mm_vision_tower, mm_mlp_adapter, and mm_language_model (not LoRA)
  • 1 epoch / 781 optimizer steps
  • Global batch size 32 on 16 H200 GPUs
  • Peak learning rate 2e-6, cosine schedule, warmup ratio 0.03
  • BF16, DeepSpeed ZeRO-3, gradient checkpointing
  • 32 video frames, maximum sequence length 32,768
  • Final training loss: 0.12651

Evaluation

All evaluations used 32 frames. Scores below are the archived January 2026 results from the same evaluation stack for the base and fine-tuned models.

Benchmark Base LV-72B SIMS-V 3Q-25K
VSI-Bench 41.18 45.03
VSI-Bench (debiased) 36.78 39.99
MME-RealWorld Lite 35.96 43.41
VideoMME 68.78 69.04
EgoSchema 66.67 63.25
OpenEQA 43.81 42.51

Usage and limitations

Load this checkpoint with the LLaVA-NeXT llava_qwen video-model code path and the qwen_1_5 conversation template. This is a research checkpoint, not a general-purpose production model. It was tuned on synthetic spatial QA and can regress on unrelated video-understanding benchmarks, as the table above shows.

The base model and this derivative use the Apache 2.0 license. Users are also responsible for the terms of the Qwen2, SigLIP, LLaVA, and SIMS-V components.

Downloads last month
13
Safetensors
Model size
73B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for spatial-training/llava-video-72b-sims-3q-25k

Finetuned
(1)
this model

Dataset used to train spatial-training/llava-video-72b-sims-3q-25k