view article Article LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge LiquidAI • 3 days ago • 35
view article Article Meta is back with Muse Glimmer: local, agentic, multimodal, and open source +2 pcuenq, merve, burtenshaw, ariG23498 • 5 days ago • 93
view article Article Making Knowledge Distillation Cheap Enough to Run at Scale MultiverseComputingCAI • 5 days ago • 29
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 17 days ago • 139
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 16 days ago • 302
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 19 days ago • 37
view article Article Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community ResterChed • 29 days ago • 195
view article Article The OlmoEarth Platform: Geospatial inference at planetary scale allenai • 17 days ago • 40
view article Article NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics nvidia • 19 days ago • 72
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published 25 days ago • 77
view article Article What building Shippy taught us about building agents allenai • about 1 month ago • 21
view article Article Fine-tune video and image models at scale with NVIDIA NeMo Automodel and 🤗 Diffusers nvidia • 28 days ago • 82