REPA: Reproducibility Evaluation via an Autonomous Pipeline Architecture

Abstract

We present REPA, a framework to autonomously reproduce text classification experiments from paper descriptions, without access to reference code or repositories, through a pipeline that runs autonomously after a single human review checkpoint on the extracted configuration. Unlike prior AI Scientist systems, REPA targets the reproduction problem directly by deconstructing it as a four-stage process incorporating protocol extraction, input preparation, experiment generation, and evaluation. On a favorable set of ten well-documented text classification papers, REPA replicates eight studies when instantiated with a GPT-4o backend, and five with Qwen3-Coder-30B. Compared to a replication rate of zero via direct prompting, our results with REPA establish a performance ceiling for current LLM automation as of mid-2026 and the importance of template scaffolding in scientific reproduction success.

Publication
ICML 2026 AI for Science Workshop
Nura Tamton
Nura Tamton
Masters Alumnus (Aug ‘25). Thesis: AI Scientist as a Reproducibility Auditor.

Master Student January 2025 Intake

Yisong Miao
Yisong Miao
Doctoral Student (Jan ‘21)

PhD Candidate January 2021 Intake

Min-Yen Kan
Min-Yen Kan
Associate Professor

WING lead; interests include Digital Libraries, Information Retrieval and Natural Language Processing.