Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Abstract

This paper enhances LLM reasoning through an iterative preference learning process inspired by AlphaZero. Monte Carlo Tree Search collects preference data iteratively, using its look-ahead ability to break instance-level rewards into step-level signals, while outcome validation and stepwise self-evaluation keep intermediate steps consistent, and Direct Preference Optimization updates the policy on the resulting step-level preferences. Theoretical analysis shows the importance of on-policy sampled data, and the method lifts Mistral-7B accuracy to 81.8% on GSM8K, 34.7% on MATH, and 76.4% on ARC-C.

Publication
NeurIPS 2024 Workshop on System-2 Reasoning at Scale
Yuxi Xie
Yuxi Xie
Doctoral Alumnus (May ‘26). Thesis: Closed-Loop Scaling: Autonomous Improvement of LLM and LVLM Reasoning

PhD Candidate January 2021 Intake

Min-Yen Kan
Min-Yen Kan
Associate Professor

WING lead; interests include Digital Libraries, Information Retrieval and Natural Language Processing.

Kenji Kawaguchi
Kenji Kawaguchi
Research Collaborator

NUS Presidential Young Professor in the Department of Computer Science, leading the Deep Learning Lab

Michael Qizhe Shieh
Michael Qizhe Shieh
Research Collaborator

Assistant Professor in the Department of Computer Science, National University of Singapore