
We propose Text-Printed Image (TPI), which generates synthetic images by directly rendering textual descriptions on a plain white canvas, bridging the modality gap between text and images for cost-efficient training of large vision-language models. TPI outperforms diffusion-based synthetic images across multiple benchmarks while preserving semantic accuracy.