Running Repro - QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture ๐ฏ Explore logs and collaborate with an AI coding agent
Running Repro: Chain-of-Thought Reasoning In The Wild Is Not Always Faithful ๐ฏ Explore and collaborate on project logbook with your AI agent
Running Repro: DFlash - Block Diffusion for Flash Speculative Decoding ๐ฏ Browse and share a compact logbook with an AI agent
Running Repro: Ambient Dataloops: Generative Models for Dataset Refinement ๐ฏ View and sync dataset refinement logs with an AI agent
Running Repro: Q-Sched: Pushing the Boundaries of Few-Step Diffusion Models with Quantization Scheduling ๐ฏ Explore project logbook and collaborate with AI agent
Running Repro: Unifying Masked Diffusion Models with Various Generation Orders and Beyond ๐ฏ Collaborate with an AI agent using a shared logbook
Running Repro: AI Cartography (ICML 2026, Vr4YScS9DT) ๐ฏ Explore AI experiment logs and sync with a coding agent
Running Repro: BioProBench (ICML 2026, 8EZamwNqld) ๐ฏ Explore and edit experimental logs in a collaborative web notebook
Running Repro: CapBencher - Give Your LLM Benchmark a Built-in Alarm ๐ Track LLM benchmarks and set automatic alerts
Running Repro: Hugging Carbon (ICML 2026, b8UehNjsF3) ๐ฏ Explore and collaborate on experiment logs in a web UI
Running Repro: SWE-rebench V2 (ICML 2026, UCAda9kS57) ๐ฏ Explore experiment logs and sync them with an AI agent
Running Repro: Conditional Equivalence of DPO and RLHF โ Explore and collaborate on project logs in a web interface
Running Repro: High-accuracy sampling for diffusion models and log-concave distributions ๐ฏ Browse experiment logs and collaborate with an AI agent
Running Repro: Benchmarking at the Edge of Comprehension ๐ฏ Explore benchmark logs and sync findings with your coding agent
Running Repro: To Grok Grokking: Provable Grokking in Ridge Regression ๐งฎ View and collaborate on a coding logbook online