view article Article BenchMIRT: What are LLM benchmarks actually measuring? allenai • 17 days ago • 22
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 15 days ago • 123
view article Article NeoMME: an efficient Multimodal-native and Multilingual Encoder Hcompany • 15 days ago • 105
view article Article Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI nico-martin, Xenova • 17 days ago • 71
view article Article State of Open Models: Summer 2026 Observations +1 AdinaY, multimodalart, irenesolaiman • Aug 14 • 203