Score
Designs, builds, or evaluates algorithms that identify and separate screen-space (2D image-plane) artifacts from authentic scene content and produce artifact-suppressed renderings or image sequences while preserving legitimate overlay appearance and consistency. This includes methods for artifact disentanglement, suppression, mitigation, and removal that may operate without ground-truth supervision or explicit artifact masks.
To address the poor generalization of conventional single-degradation image restoration methods under realistic scenarios where multiple degradations (e.g., noise, blur, weather artifacts) co-occur, this paper proposes a unified All-in-One Image Restoration (AiOIR) paradigm. We first establish a systematic taxonomy for AiOIR; then introduce a three-dimensional evaluation framework covering prior modeling, generalization capability, and learning paradigms; and finally design an adaptive architecture integrating multi-task learning, meta-learning, degradation-aware attention, and a shared-specialized dual-path network. Extensive benchmarking is conducted on mainstream datasets (Rain13k, RealBlur, DPSR) using PSNR, SSIM, and LPIPS metrics, with open-source method comparisons. Contributions include: (i) the first structured AiOIR survey, (ii) an objective performance benchmark, (iii) an open-source codebase (GitHub), and (iv) identified future directions—scalable architectures, disentangled representations, and dynamic inference.
本文通过引入包含合成与真实视频对的BeyondMasks基准和CORE评估协议,解决视频对象移除中的因果一致性问题。
This work addresses the challenge that screen-space artifacts in real-world images—such as lens smudges and UI watermarks—are often erroneously reconstructed as floating 3D objects, severely degrading novel view synthesis quality. To resolve this, we propose the first unsupervised framework that jointly optimizes a 3D Gaussian point cloud and a learnable 2D overlay layer. By leveraging multi-view geometric consistency, our method automatically disentangles static artifacts from genuine 3D structure without requiring manual annotations or prior knowledge. Evaluated on both synthetic and real-world datasets, the approach achieves up to a 9 dB PSNR improvement over the original 3D Gaussian Splatting baseline, significantly enhancing reconstruction fidelity while accurately preserving artifact content.
This work addresses the geometric and photometric degradation in 3D Gaussian Splatting (3DGS) under sparse-view settings by proposing a unified inpainting framework leveraging video diffusion models. The authors construct a large-scale video dataset comprising 107.5K paired samples and introduce an isomorphic dual-model architecture featuring fine-grained 3DGS artifact classification and an Artifact-Aware Triplet Fusion mechanism guided by artifact heatmaps. For the first time, intensity-aware restoration is integrated into the self-attention structure, enabling precise spatiotemporal consistent inpainting. The proposed method significantly outperforms existing approaches in sparse novel-view synthesis and robust 3D reconstruction, effectively enhancing multi-view consistency and generalization capability.
Addressing blind image restoration under unknown real-world degradations without ground-truth images, this work proposes a holistic solution. First, we introduce a learnable degradation chain estimator to accurately model complex, realistic degradations. Second, we design a consistency-driven, plug-and-play diffusion prior framework enabling end-to-end lightweight optimization. Third, we pioneer reference-free proxy metrics—MSE and LPIPS computed on synthetically degraded samples—that overcome the longstanding challenge of unreliable performance evaluation in blind restoration. To our knowledge, this is the first work unifying degradation modeling, restoration algorithm design, and no-reference assessment within a single coherent pipeline. Extensive experiments demonstrate substantial improvements in ranking accuracy over SOTA methods across multiple blind restoration benchmarks, significantly enhancing algorithmic assessability, comparability, and practical applicability.
Evaluating 3D reconstruction quality without ground-truth images faces two key challenges: the absence of reliable reference views and the inability of prevailing no-reference metrics to localize fine-grained artifacts in novel synthesized views. To address this, we propose Puzzle Similarity (PS), the first no-reference metric capable of spatially localizing artifact regions. Its core innovation lies in modeling scene-specific distributions via local patch statistics extracted from the input multi-view images, enabling adaptive distribution matching for fine-grained, reference-free artifact localization. Evaluated on human-perception datasets, PS significantly outperforms existing full-reference and no-reference metrics, achieving high correlation with subjective quality judgments. Crucially, PS operates entirely without ground-truth imagery, making it suitable for downstream applications such as artifact-driven automatic restoration, acquisition optimization, and sparse-view reconstruction.
Existing image inpainting methods often struggle to preserve fine details and produce natural-looking results when handling complex semi-transparent degradations—such as cracks and stains—in artistic images, primarily due to their reliance on predefined degradation models. This work addresses this limitation by introducing the first controllable benchmark for blind restoration of artistic imagery, accompanied by a new dataset, MDTD-Art, which features semi-transparent degradation masks at multiple opacity levels. Through systematic evaluation of general-purpose inpainting models, image editing models, and vision-language models, the study demonstrates that image editing models enhanced with structured prompt engineering consistently outperform specialized inpainting approaches across arbitrary degradations. Notably, these models excel when guided by semantic cues that emphasize structural coherence and fine-grained detail, underscoring the critical role of semantically controllable generation in blind restoration of artistic images.
该研究针对3D高斯点绘在未观察区域的重建问题,提出了一种新的框架,通过独立视图检测机制和质量感知掩模模块提高场景外推的质量。
This work addresses the growing challenge of verifying the authenticity of increasingly realistic AI-generated images, a task for which existing methods struggle to jointly perform detection and artifact correction. The paper proposes GenShield, the first unified autoregressive framework that integrates detection and controllable artifact repair through a synergistic “diagnose-and-repair” mechanism operating in a closed loop, enabling interpretable detection and targeted restoration. GenShield innovatively establishes a mutually reinforcing relationship between detection and repair, incorporating a Visual Chain-of-Thought curriculum learning strategy that supports multi-step self-explanatory refinement with an explicit stopping criterion. Evaluated on both established detection benchmarks and a newly curated large-scale dataset of artifact–repair image pairs using a comprehensive protocol, GenShield achieves state-of-the-art performance and demonstrates exceptional generalization across diverse generation models.
Existing approaches struggle to jointly handle scene text editing tasks—deletion, generation, and replacement—within a unified framework that simultaneously ensures precise textual appearance control and background integrity. To address this, this work proposes a unified model that decomposes complex text editing into two atomic operations: rendering and erasure. It introduces Overlay-Reference Positional Encoding (ORPE) to achieve pixel-level layout fidelity and exemplar-driven style control, complemented by a Region-Adaptive Suppression (RAS) strategy to ensure clean text removal. The study also establishes TextWand-Bench, the first comprehensive benchmark for general scene text editing. Experimental results demonstrate that the proposed method significantly outperforms both open-source and closed-source models across all three editing tasks in terms of text accuracy, layout-style consistency, and overall image quality.
本文提出一种无需训练的零样本图像篡改定位方法,通过分析单张可疑图片的噪声残差模式,评估图像真实性。