HumanCLAW: Can Vision-Language Models Act Through a Body? Paper • 2607.27180 • Published 5 days ago • 71
IAM: Identity-Aware Human Motion and Shape Joint Generation Paper • 2604.25164 • Published Apr 28 • 2
UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors Paper • 2603.15975 • Published Mar 16 • 3
LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens Paper • 2602.12370 • Published Feb 12
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding Paper • 2504.13915 • Published Apr 10, 2025