Research arXiv cs.AI

RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?

reproducibilitybenchmarkAI agentsNeurIPS

Reproducing a machine learning paper involves most research steps, from installing software and debugging to running experiments—work that AI agents increasingly do. RECLAIM is introduced as a benchmark of 100 NeurIPS 2025 papers that can be rebuilt yearly from new conferences. For each paper, the result to reproduce, what counts as a successful reproduction, and a GPU-hour budget are fixed in advance, and an agent must reproduce that result using the paper and whatever its authors released.

What the authors released decides the difficulty tier. Run-tier releases include code, data, and weights; Retrain-tier releases lack weights, so the agent must train the model itself; Reimplement-tier releases lack code, so the agent must write it. A separate language model grades runs from logs and outputs rather than relying on the agents' reports.

Four agents were run once per paper. The best agent in each tier reproduces only 41% of Run-tier papers, 27% at Retrain, and 15% at Reimplement, where every agent does worst. Failed attempts use on average 29% of their budget, so most stop with budget left. The most common agent error is writing the method without checking any part against the paper's numbers, occurring in 63 of 400 runs.

Read original →

← Back to home