Video understanding, multimodal large language models, reinforcement learning, temporal and spatial grounding, tracking, segmentation, and spatial intelligence.