Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development Paper • 2608.13417 • Published 5 days ago • 45
view article Article TutorMoments: Do AI tutors know when to help and when to hold back? allenai • 11 days ago • 30
Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements Paper • 2607.28661 • Published 27 days ago • 14
When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings Paper • 2608.03994 • Published 14 days ago • 8
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published 18 days ago • 96
Instella-MoE ✨ Collection Family of fully open 16B MoE LLM with 2.8B active params per token, trained on AMD Instinct™ MI300 & MI325 GPUs. • 6 items • Updated 21 days ago • 18
microsoft/VibeVoice-ASR-BitNet Automatic Speech Recognition • 0.3B • Updated 25 days ago • 16.6k • 182
view article Article NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval nvidia • Jul 16 • 59
Nemotron 3 Embed Collection Open embedding models for enterprise RAG, agentic retrieval, code search, and agent memory. • 3 items • Updated 7 days ago • 36
view article Article Distillation in 2026 (so far): which frontier models use it and how sergiopaniego • Jul 8 • 21