Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models

Abstract

We propose Text-Printed Image (TPI), which generates synthetic images by directly rendering textual descriptions on a plain white canvas, bridging the modality gap between text and images for cost-efficient training of large vision-language models. TPI outperforms diffusion-based synthetic images across multiple benchmarks while preserving semantic accuracy.

Publication
CVPR 2026
Futa Waseda
Futa Waseda
Project Assistant Professor | Trustworthy AI: Robustness, Reliability, and VLM Defense

Related