
Large vision-language models (LVLMs) are vulnerable to typographic attacks—misleading text within images that overrides visual understanding. We introduce Read-or-Ignore VQA (RIO-VQA) and its benchmark RIO-Bench, a unified evaluation framework assessing when models should read text versus ignore it based on context. Our findings show that strong LVLMs and existing defenses fail to balance typographic robustness and text-reading capability. We propose a data-driven defense mechanism enabling adaptive selective text use, moving beyond previous non-adaptive approaches.
TL;DR: VLMs must sometimes read text in images and sometimes ignore it — but no model does both well. We propose RIO-Bench, the first benchmark unifying typographic-attack robustness and text recognition evaluation. We further introduce a data-driven defense that adaptively selects when to use in-image text, outperforming prior non-adaptive approaches.