Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space Paper • 2510.04476 • Published Oct 6, 2025 • 16
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design Paper • 2511.17127 • Published Nov 21, 2025 • 4
The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility Paper • 2605.19537 • Published May 20 • 2
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression Paper • 2510.13999 • Published Oct 15, 2025 • 23