Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs Paper • 2609.10355 • Published 6 days ago • 14
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing Paper • 2608.06146 • Published Aug 6 • 24
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Paper • 2607.25948 • Published Jul 28 • 20
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published Jul 29 • 141