Super-Resolution in The Right Latent Space: A Frozen Vision-Foundation Substrate

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the mapping difficulties and severe hallucination artifacts in traditional super-resolution methods caused by ill-structured latent spaces. We propose a novel super-resolution paradigm built upon the latent space of vision foundation models. Specifically, a frozen DINOv3-L fusion layer serves as an ideal latent substrate, leveraging its semantic hierarchy to guide faithful reconstruction. A lightweight decoder is optimized via a joint adversarial and reconstruction training strategy, enabling precise mapping of degraded image embeddings onto the latent manifold and single-step pixel reconstruction. Experiments demonstrate that this approach achieves state-of-the-art fidelity-perception trade-offs across multiple benchmarks, significantly outperforming VAE-based latent methods while requiring only 37 milliseconds for inference on 512×512 images.
📝 Abstract
In an image latent space, the embeddings of high-resolution, natural, and sharp images form a manifold. Degradation of high-resolution images pushes their embeddings off this manifold. Real-world super-resolution (SR) then becomes the task of mapping the degraded embedding back onto this manifold --- not anywhere on the manifold, but to the point that preserves what the input still carries, both its semantics and pixel details. Every published method implements this mapping in a reconstruction-oriented latent space or pixel space. We claim these spaces are the wrong substrates for SR. Low-resolution and degraded images are embedded far from the manifold, making the mapping difficult and expensive. The lack of semantic information in these substrates also makes it difficult to navigate to the faithful point on the manifold, causing severe hallucination when degradation is heavy. Thus, restoring in a suitable latent space is crucial to the SR task. We show that the latent space of 23 fused layers of a frozen DINOv3-L is one such space that makes the SR task easier. Degraded images are embedded near the manifold. Moreover, this substrate contains a hierarchy of information, from pixel record to degradation robust semantics, guiding the model to find the faithful point on the manifold. On this substrate, a 415M decoder is trained under reconstruction and adversarial objectives to map the degraded embeddings back and decode to pixel space in one pass. The resulting model, RAESR, attains the best fidelity--perception trade-off among state-of-the-art adversarial and diffusion-based restorers on RealSR, DRealSR, LSDIR and DIV2K-Val, at 37 ms per 512 by 512 image on a single H20 GPU. Swapping the substrate for a VAE latent under an identical recipe loses on every metric.
Problem

Research questions and friction points this paper is trying to address.

Super-Resolution
Latent Space
Image Manifold
Hallucination
Degradation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Super-Resolution
Latent Space
Vision Foundation Model
Frozen DINOv3
Fidelity-Perception Trade-off
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wanzhou Lei
University of California, Berkeley
C
Cuifeng Sheng
Alibaba Group
Y
Yanjin He
University of Michigan, Ann Arbor
Maohua Li
Maohua Li
Hohai University
Spiking Neural Networks
H
Hua Yuan
Southeast University
Per-Olof Persson
Per-Olof Persson
University of California, Berkeley
Computational fluid and solid mechanicshigh-order discontinuous Galerkin methodsfluid-structure interactionunstructured me
H
Hanlin Tang
Alibaba Group