Humans perceive images at multiple levels, from low-level object recognition to high-level semantic interpretation such as behavior understanding, and subtle low-level differences can flip the high-level reading of a scene. MVP-Bench is the first vision-language benchmark that systematically evaluates both low- and high-level visual perception, constructed across natural and synthetic images to test how manipulated content influences model perception. Diagnosing 10 open-source and 2 closed-source LVLMs shows high-level perception remains challenging: GPT-4o reaches 56% accuracy on Yes/No questions against 74% in low-level scenarios.