π€ AI Summary
This study addresses the high computational cost of multi-step inference and semantic distortion caused by erroneous text priors in diffusion-based text image super-resolution. To this end, we propose a single-step latent adaptation framework that eliminates iterative image-text diffusion, achieving latent space adaptation in a single step for the first time. By incorporating confidence-weighted conditioning and a lightweight residual correction module, the proposed method effectively suppresses OCR error accumulation while precisely recovering stroke-level details. Extensive experiments on the CTR-TSR-Test and RealCE-200 benchmarks demonstrate state-of-the-art performance, yielding PSNR improvements of at least 2.72 dB over existing approaches. Overall, this work enables efficient and high-fidelity reconstruction of degraded text images, offering a practical solution to the limitations of current diffusion-based methods.
π Abstract
Text image super-resolution (TSR) aims to recover visually faithful and readable text under unknown degradations. Existing diffusion-based methods typically rely on multi-step prediction of either the high-resolution image or its text prior, resulting in prohibitive computational cost and inference latency. More critically, an erroneous text prior may be repeatedly injected into the denoising process, causing image and text predictions to reinforce each other and progressively amplify an early recognition error into a sharp yet semantically incorrect character. To address these limitations, we propose TOLA, a Text-aware One-step Latent Adaptation framework without iterative image-text diffusion. TOLA consists of two key modules. First, a confidence-weighted text conditioning module constructs the semantic condition only once and suppresses unreliable OCR predictions before they contaminate image reconstruction. Second, a lightweight latent residual correction module explicitly estimates and corrects the structured residual errors to recover missing or distorted stroke details. Extensive experiments demonstrate our state-of-the-art performance across all evaluation metrics on both CTR-TSR-Test ($\times 4$) and RealCE-200 benchmarks. It is worth noting that our TOLA consistently surpasses existing diffusion-based TSR methods by at least 2.72 dB in PSNR on CTR-TSR-Test.