ComicVQA: A Benchmark for Visual Reasoning in Multimodal LLMs

Abstract

ComicVQA is a comics-based benchmark for multimodal LLMs, pairing a task on fine-grained visual grounding in panels with a task on sequential narrative structure. Proprietary models reach 62.6% and 46.4% on the two tasks and open-source models reach 47.7% and 26.9%, while human annotators exceed 83%, indicating that current models rely on temporal cues instead of detailed visual reasoning.

Publication
Findings of the Association for Computational Linguistics: ACL 2026
Esther Gan
Esther Gan
Doctoral Student (Aug ‘23)
Co-Supervised by Michael Shieh

PhD Candidate August 2023 Intake

Kenji Kawaguchi
Kenji Kawaguchi
Research Collaborator

NUS Presidential Young Professor in the Department of Computer Science, leading the Deep Learning Lab

Min-Yen Kan
Min-Yen Kan
Associate Professor

WING lead; interests include Digital Libraries, Information Retrieval and Natural Language Processing.

Michael Qizhe Shieh
Michael Qizhe Shieh
Research Collaborator

Assistant Professor in the Department of Computer Science, National University of Singapore