SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers

Published in COLM 2025, 2025

We introduce SciReplicate-Bench, the first benchmark targeting the reliable reproduction of algorithmic results described in scientific papers with agentic LLM pipelines. The benchmark covers memory management, tool grounding, and execution tracking, enabling systematic evaluation of agent behaviors. Project resources are available on the website.

Recommended citation: Yanzheng Xiang, Hanqi Yan, Shuyin Ouyang, Lin Gui, Yulan He. 2025. "SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers." In COLM 2025.
Download Paper

Code