Community Hugging Face Blog

What We Learned by Reproducing 2,200 papers from ICML

ICML 2026reproducibilitycoding agentshackathon

The hackathon was motivated by longstanding questions about how reproducible AI research is, questions now exacerbated by scale. ICML 2026 received 23,918 submissions and accepted 6,352 papers, roughly double the previous year, a trend partly driven by AI agents making it faster to run experiments and write them up. Yet reviewing capacity has not grown accordingly; reviewers are often volunteers who lack time or expertise. One accepted ICML 2026 spotlight review admitted: "My low confidence score is because I did not check all the proofs carefully."

From July 15 to August 2, 2026, the ICML 2026 Open Reproductions challenge invited the community to pick a paper from the 6,341 accepted papers, which were indexed with abstracts and core scientific claims extracted. Participants used coding agents like Claude Code, Codex, Cursor, and OpenResearch's orx, and were encouraged to reproduce the same paper multiple times. A streamlined interface let agents pull the paper, its claims, and the challenge details. In 19 days, more than 1,200 participants published 6,816 Trackio logbooks reproducing 2,226 papers — about a third of the conference.

The post reports on reproductions done well, falsifications and what happened when they checked them, and conversations with authors. The key takeaway is that coding agents can now read a paper, write code, launch experiments, and report findings in an afternoon, in parallel, thousands of times over — whereas checking a paper used to cost a reviewer a weekend. The findings suggest a significant shift in the division of labor between humans and AI agents in research validation, with humans potentially focusing on interpretation and judgment rather than manual verification.

Read original →

← Back to home