Self-Evaluation Guided Beam Search for Reasoning

Abstract

Breaking a problem into intermediate steps helps LLM reasoning, but longer reasoning chains accumulate uncertainty and error, making accurate final answers hard to elicit. This paper introduces a stepwise self-evaluation mechanism that guides and calibrates the reasoning process, integrated into decoding through stochastic beam search that balances exploitation and exploration with temperature-controlled randomness. The approach improves few-shot accuracy over the corresponding Codex-backboned baselines by 6.34%, 9.56%, and 5.46% on GSM8K, AQuA, and StrategyQA.

Publication
Advances in Neural Information Processing Systems 36 (NeurIPS 2023)
Yuxi Xie
Yuxi Xie
Doctoral Alumnus (May ‘26). Thesis: Closed-Loop Scaling: Autonomous Improvement of LLM and LVLM Reasoning

PhD Candidate January 2021 Intake

Kenji Kawaguchi
Kenji Kawaguchi
Research Collaborator

NUS Presidential Young Professor in the Department of Computer Science, leading the Deep Learning Lab

Min-Yen Kan
Min-Yen Kan
Associate Professor

WING lead; interests include Digital Libraries, Information Retrieval and Natural Language Processing.

Michael Qizhe Shieh
Michael Qizhe Shieh
Research Collaborator

Assistant Professor in the Department of Computer Science, National University of Singapore