Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models

Abstract

Large vision-language models (LVLMs) are vulnerable to typographic attacks—misleading text within images that overrides visual understanding. We introduce Read-or-Ignore VQA (RIO-VQA) and its benchmark RIO-Bench, a unified evaluation framework assessing when models should read text versus ignore it based on context. Our findings show that strong LVLMs and existing defenses fail to balance typographic robustness and text-reading capability. We propose a data-driven defense mechanism enabling adaptive selective text use, moving beyond previous non-adaptive approaches.

Publication
ECCV 2026

TL;DR: VLMs must sometimes read text in images and sometimes ignore it — but no model does both well. We propose RIO-Bench, the first benchmark unifying typographic-attack robustness and text recognition evaluation. We further introduce a data-driven defense that adaptively selects when to use in-image text, outperforming prior non-adaptive approaches.

Futa Waseda
Futa Waseda
Project Assistant Professor | Trustworthy AI: Robustness, Reliability, and VLM Defense

Related