VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing Paper • 2603.29852 • Published Feb 22 • 6
Mem-$π$: Adaptive Memory through Learning When and What to Generate Paper • 2605.21463 • Published May 20 • 9
Toward Open Weight Models Without Risks: Separating Public and Private Capabilities in LLMs Paper • 2606.21638 • Published 21 days ago • 7
PrivacyAlign: Contextual Privacy Alignment for LLM Agents Paper • 2606.21710 • Published 21 days ago • 3
PrivacyAlign: Contextual Privacy Alignment for LLM Agents Paper • 2606.21710 • Published 21 days ago • 3
view article Article MosaicLeaks: Can your research agent keep a secret? ServiceNow • 22 days ago • 13
SproutRAG: Attention-Guided Tree Search with Progressive Embeddings for Long-Document RAG Paper • 2606.18381 • Published 24 days ago • 19
MCompassRAG: Topic Metadata as a Semantic Compass for Paragraph-Level Retrieval Paper • 2606.18508 • Published 24 days ago • 21
Structured Distillation of Web Agent Capabilities Enables Generalization Paper • 2604.07776 • Published Apr 9 • 23
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents Paper • 2603.24440 • Published Mar 25 • 99
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents Paper • 2603.24440 • Published Mar 25 • 99
LLM2Vec-Gen: Generative Embeddings from Large Language Models Paper • 2603.10913 • Published Mar 11 • 44
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA Paper • 2505.16293 • Published May 22, 2025 • 3
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation Paper • 2508.16763 • Published Aug 22, 2025 • 2