MVP-Bench: Can Large Vision-Language Models Conduct Multi-level Visual Perception Like Humans?

Abstract

Humans perceive images at multiple levels, from low-level object recognition to high-level semantic interpretation such as behavior understanding, and subtle low-level differences can flip the high-level reading of a scene. MVP-Bench is the first vision-language benchmark that systematically evaluates both low- and high-level visual perception, constructed across natural and synthetic images to test how manipulated content influences model perception. Diagnosing 10 open-source and 2 closed-source LVLMs shows high-level perception remains challenging: GPT-4o reaches 56% accuracy on Yes/No questions against 74% in low-level scenarios.

Publication
Findings of the Association for Computational Linguistics: EMNLP 2024
Guanzhen Li
Masters Alumnus (Aug ‘23). Thesis: Edited Media Understanding

Graduate Student August 2023 Intake

Yuxi Xie
Yuxi Xie
Doctoral Alumnus (May ‘26). Thesis: Closed-Loop Scaling: Autonomous Improvement of LLM and LVLM Reasoning

PhD Candidate January 2021 Intake

Min-Yen Kan
Min-Yen Kan
Associate Professor

WING lead; interests include Digital Libraries, Information Retrieval and Natural Language Processing.