Batch-Transformer for Scene Text Image Super Resolution

preprint OA: closed CC-BY-4.0
📄 Open PDF View at publisher

Abstract

Abstract Recognizing low-resolution text images is challenging as they often lose their detailed information, leading to poor recognition accuracy. Moreover, the traditional methods, based on deep Convolutional Neural Networks (CNNs), are not effective enough for some low-resolution text images with dense characters. In this paper, a novel CNN-based Batch-Transformer network for Scene Text Image Super Resolution (BT-STISR) method is proposed to address this problem. In order to obtain the text information for text reconstruction, a pre-trained text prior module is employed to extract text information. Then a novel two pipeline batch transformer-based module is proposed, leveraging self-attention and global attention mechanisms to exert the guidance of text prior to the text reconstruction process. Experimental study on a benchmark dataset TextZoom shows that the proposed method BT-STISR achieves the best state-of-the-art performance in terms of Structural Similarity (SSIM) and Peak signal-to-noise ratio (PSNR) metrics compared to some latest methods.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0