V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization

Abstract

Large vision-language models hallucinate in part because they over-rely on the language model backbone, which introduces bias from language priors and leaves insufficient attention on the visual input. This paper mitigates that over-reliance through preference learning, proposing Vision-guided Direct Preference Optimization (V-DPO) to strengthen visual context learning at training time. Evaluated on a synthetic dataset containing both response-contrast and image-contrast preference pairs, V-DPO improves over baseline methods across hallucination benchmarks and gains the most from image-contrast data.

Publication
Findings of the Association for Computational Linguistics: EMNLP 2024
Yuxi Xie
Yuxi Xie
Doctoral Alumnus (May ‘26). Thesis: Closed-Loop Scaling: Autonomous Improvement of LLM and LVLM Reasoning

PhD Candidate January 2021 Intake

Guanzhen Li
Masters Alumnus (Aug ‘23). Thesis: Edited Media Understanding

Graduate Student August 2023 Intake

Xiao Xu
Xiao Xu
CSC Postgraduate Visiting Student (Sep ‘23) Project: Vision Language Models

Visiting student; interests include Vision-Language Learning, Large Language Models.

Min-Yen Kan
Min-Yen Kan
Associate Professor

WING lead; interests include Digital Libraries, Information Retrieval and Natural Language Processing.