Score
Designs and implements rendering algorithms that synthesize spatially‑varying optical defocus by applying depth‑dependent blur to image or scene representations using depth maps or buffers to determine local blur amount. Builds occlusion‑aware compositing and visibility rules to correctly handle depth boundaries and compose blurred foreground and background regions with the proper occlusion ordering.
Fixed-focus, large-aperture cameras (e.g., smart glasses) suffer from depth-dependent defocus blur and spatially varying optical aberrations due to shallow depth of field; existing deblurring models trained on open-source datasets exhibit significant domain shift and poor generalization to such real-world imaging conditions. Method: We propose a physics-driven synthetic image generation framework that jointly models depth-dependent defocus blur and spatially varying aberrations for the first time—enabling simulation-to-real domain alignment without fine-tuning on real data. A lightweight low-resolution synthetic dataset is used to train an end-to-end deep network capable of restoring high-resolution (12 MP) real images. Contribution/Results: Our method achieves state-of-the-art performance across diverse challenging scenarios, demonstrating superior efficiency, scalability, and practical deployability compared to prior approaches.
This work addresses the highly complex optimization problem of recovering depth from multiple defocused images by proposing an efficient global optimization method based on alternating minimization. Specifically, when the depth map is fixed, the all-in-focus image is optimized via a linear formulation; conversely, with the all-in-focus image held constant, the depth at each pixel is solved independently and in parallel. This approach is the first to enable efficient global optimization for high-resolution depth estimation without relying on deep learning models. By integrating convex optimization with parallel grid search and directly leveraging the physical optics-based forward model, the method achieves significant performance gains over existing techniques on both synthetic and real defocus datasets, supporting accurate reconstruction of high-resolution depth maps.
Existing defocus deblurring datasets struggle to support model generalization across diverse cameras and lenses due to limited optical diversity and insufficient realism. To address this, this work proposes the first scalable synthetic defocus dataset generation framework that accommodates heterogeneous compound lenses, occlusion handling, and realistic imaging pipelines. Operating in radiometrically linear space, the framework integrates Debye CZT wave-optics point spread functions, depth-aware blur rendering, and camera ISP simulation to produce the high-fidelity synthetic dataset CLDefocus. Experiments demonstrate that models trained on CLDefocus significantly outperform current methods in cross-device defocus deblurring tasks. Furthermore, the study reveals a systematic bias in full-reference evaluations when using real-captured data, underscoring the advantages of the proposed physically grounded synthesis approach.
Existing image generation methods struggle to precisely control occlusion relationships among objects: text-guided approaches lack geometric consistency, while layout-to-image methods neglect depth-order modeling. This paper introduces the first occlusion-control framework that operates in the latent space of pretrained diffusion models without fine-tuning, incorporating volumetric rendering principles. By jointly estimating object transmittance and density—and integrating occlusion-aware latent-space manipulations—it enables physically consistent foreground-background reasoning. The method supports fine-grained control over multiple visual attributes, including transparency, fog concentration, and illumination intensity. Evaluated on multiple benchmarks, it achieves significantly higher occlusion accuracy than state-of-the-art methods, while relying solely on off-the-shelf pretrained diffusion models and requiring no additional training.
Existing 3D Gaussian Splatting (3DGS) methods rely on the pinhole camera model, which cannot represent defocus blur—limiting applications such as post-capture refocusing, all-in-focus rendering, and deblurring. This work proposes the first differentiable depth-of-field rendering framework for finite-aperture cameras, enabling explicit modeling of adjustable aperture and focus distance within 3DGS for the first time. It supports 3D reconstruction and post-capture refocusing from multi-view blurry images. A novel learnable Circle-of-Confusion (CoC)-aware optimization mechanism is introduced to identify in-focus regions and jointly optimize geometric and blur cues. The method requires no camera calibration. Experiments demonstrate significant improvements in refocusing accuracy, controllability of bokeh effects, and all-in-focus image synthesis quality, outperforming state-of-the-art 3DGS approaches across multiple tasks.
Existing single-image refocusing methods rely on all-in-focus inputs, synthetic data, and exhibit limited aperture control. This work proposes the first end-to-end depth-of-field refocusing framework capable of handling arbitrarily blurred inputs, enabling simultaneous all-in-focus restoration and controllable, photorealistic bokeh synthesis. Methodologically, we introduce a semi-supervised training paradigm that jointly leverages synthetic paired data and unpaired real-world bokeh images; incorporate EXIF metadata-driven, physics-aware modeling to explicitly encode lens optical properties; and integrate three core modules—DeblurNet (for deblurring), BokehNet (for parametric bokeh generation), and a text-guided diffusion-based control module—to support continuous aperture adjustment and non-circular aperture blur. Our approach achieves state-of-the-art performance across three major benchmarks—deblurring, bokeh synthesis, and refocusing—demonstrating significant improvements in real-scenario applicability and editing flexibility.
Existing 3D layout-guided image generation methods struggle to accurately model occlusion relationships among objects, often resulting in inconsistent geometry and scale for occluded entities in synthesized scenes. To address this limitation, this work proposes SeeThrough3D, the first approach to explicitly model occlusion relations in text-to-image generation. By introducing an Occlusion-aware 3D Scene Representation (OSCR), objects are represented as semi-transparent 3D bounding boxes, enabling explicit reasoning about occluded regions. The method integrates a flow-based diffusion model with a mask-based self-attention mechanism to achieve precise 3D layout control while preserving occlusion consistency. SeeThrough3D supports multi-object binding and viewpoint-consistent occlusion synthesis, generalizes well to unseen object categories, and significantly enhances both the realism and layout fidelity of complex scene generation.
This work addresses the challenge of novel view synthesis from motion-blurred inputs, a setting where existing methods struggle due to inconsistent geometry and appearance, while conventional deblurring approaches rely on costly per-scene optimization and exhibit limited generalization. We propose DeblurNVS, the first framework capable of synthesizing sharp, multi-view-consistent novel views from sparse, motion-blurred images without per-scene optimization. Our method leverages a geometric latent diffusion model to recover coherent 3D structure and correspondences from blurred inputs, which are then integrated with either neural radiance fields or 3D Gaussian splatting to render high-fidelity views. Extensive experiments on both synthetic and real-world blurred scenes demonstrate that DeblurNVS significantly outperforms state-of-the-art alternatives, producing structurally stable and visually clear results while enabling efficient and generalizable sparse-view synthesis.
To address two key challenges in high-fidelity rendering of dynamic specular scenes— inaccurate reflection direction estimation and physical model distortion—this paper proposes a residual material-enhanced 2D Gaussian splatting representation, introduces the first dynamic environment Gaussian modeling framework, and establishes a diffuse-specular decoupled hybrid rendering pipeline integrating rasterization and ray tracing. Leveraging physics-driven reflection composition and a coarse-to-fine multi-stage optimization strategy, the method ensures both interpretability and physical fidelity. Evaluated on dynamic scene benchmarks, our approach achieves state-of-the-art quantitative performance and produces qualitatively superior results: sharper, physically consistent specular highlights. Notably, it is the first method to enable high-accuracy, physically coherent 3D reconstruction and rendering of dynamic specular effects.
This work addresses the frequent logical errors in occlusion relationships—particularly within densely overlapping regions—that plague current text-to-image diffusion models. To resolve this, the authors propose a training-free, plug-and-play framework that introduces, for the first time, a depth-aware Attention Arbitration Mechanism (AAM) to modulate inter-object attention competition in accordance with realistic depth ordering. This mechanism is complemented by Spatial Compactness Control (SCC) to enhance structural coherence. Evaluated on the newly introduced OcclBench benchmark, the method demonstrates superior performance over state-of-the-art approaches in both occlusion accuracy and visual fidelity, significantly strengthening the spatial hierarchy representation capabilities of diffusion models.