Score
Designs and implements methods that generate novel views and renderings of scenes by estimating continuous-depth fields and producing images or volumetric slices at arbitrary viewpoints and intermediate depths. This includes end-to-end, single-model, training-free and zero-shot approaches that take sparse or anisotropic z-stacks or multi-view images as input and interpolate depths to produce isotropic 3D volumes and arbitrary-direction sectioning.
Motivated by the urgent demands of AR/VR and digital twin applications for fast, generalizable, and deployment-friendly 3D reconstruction and novel view synthesis, this paper presents a systematic survey of feedforward deep learning methods—covering dominant representations including point clouds, 3D Gaussian splatting, and neural radiance fields—and focuses on three key challenges: pose-free input, dynamic scene modeling, and 3D-aware content generation. We propose the first unified taxonomy tailored to the feedforward paradigm, revealing inherent trade-offs between inference efficiency and cross-scene generalization. By integrating self-supervised learning, differentiable rendering, and multimodal input strategies—and leveraging standardized evaluation protocols and large-scale benchmarks—we comprehensively assess accuracy, latency, and robustness. Our analysis provides principled guidance and empirically grounded technology selection criteria for industrial-grade 3D vision systems.
This work addresses the challenge of balancing accuracy and efficiency in real-time novel view synthesis for radiance fields. We propose Radiance Grid—a tetrahedral voxel representation with uniform density, constructed via Delaunay tetrahedralization—where radiance field parameters are defined per tetrahedral cell. Coupled with a Zip-NeRF–style backbone network, the formulation ensures field continuity under topological changes. We design a dedicated rasterizer compatible with ray tracing, enabling hardware-efficient, accurate volume rendering. Compared to existing radiance field methods, Radiance Grid achieves higher rendering speed and fidelity on consumer-grade GPUs. It supports real-time novel view synthesis, fisheye distortion modeling, physics-based simulation, interactive scene editing, and isosurface mesh extraction. Crucially, it is the first approach to achieve efficient, differentiable volume rendering while preserving geometric precision—enabling high-fidelity, real-time applications without sacrificing structural accuracy.
Addressing the challenge of real-time multi-view stereo (MVS) reconstruction and novel view synthesis (NVS) under sparse-view settings, this paper introduces the first end-to-end feedforward 2D Gaussian splatting framework. The method directly regresses generalizable 2D Gaussian parameters, jointly optimizing geometric reconstruction accuracy and rendering quality. To enhance precision, speed, and cross-dataset generalization, it incorporates multi-view feature distillation and explicit MVS supervision, integrating pre-trained visual features. Experimental results demonstrate state-of-the-art performance on DTU (Chamfer distance), significant improvements over existing methods on BlendedMVS and Tanks and Temples, and inference speed approximately 100× faster than implicit volumetric rendering. The core contribution lies in the first differentiable, generalizable, and real-time feedforward 2D Gaussian splatting reconstruction pipeline tailored for sparse-view scenarios.
Existing 3D Gaussian Splatting (3DGS) rasterization methods suffer from popping artifacts, view-dependent density, and difficulty modeling lens effects (e.g., defocus blur, fisheye distortion). This paper introduces the first differentiable volumetric rendering framework based on ellipsoidal primitives, prioritizing emission. It eliminates discontinuity artifacts and decouples density from viewing direction via analytic ellipsoid-ray intersection, physically consistent volume integration, and alpha-free end-to-end optimization. The method enables real-time GPU-accelerated ray tracing, rendering 720p frames at ~30 FPS on an RTX 4090. On large-scale scenes such as Zip-NeRF, it achieves the sharpest reconstruction quality among current real-time approaches, significantly suppressing blending artifacts while naturally supporting complex lens effects.
Existing 3D novel view synthesis (NVS) methods rely heavily on dense 3D annotations or multi-view inputs, limiting generalization; conversely, 3D-free approaches enable text-driven generation of complex scenes but lack precise camera control. This paper proposes a single-image-driven, camera-controllable NVS framework. For the first time, it incorporates a pre-trained NVS model as a weak geometric prior into a 3D-free diffusion architecture and explicitly encodes camera pose in the CLIP embedding space. Built upon Stable Diffusion, our method integrates cross-modal CLIP enhancement, knowledge distillation, and pose-conditioned embedding—requiring neither multi-view nor 3D supervision. Evaluated on diverse real-world scenes, it significantly outperforms state-of-the-art methods, enabling accurate, arbitrary-view synthesis with high fidelity, geometric consistency, and fine-grained detail. Both qualitative and quantitative evaluations demonstrate leading performance.
To address overfitting in 3D Gaussian Splatting (3DGS) under single-view supervision—leading to artifacts in novel-view synthesis and inaccurate geometric reconstruction—this paper proposes a multi-view collaborative optimization framework. Our method introduces three key innovations: (1) a novel multi-view regularization paradigm that jointly enforces consistency across multiple views; (2) an intrinsic-cross-guided coarse-to-fine training strategy integrating multi-scale geometric and appearance priors; and (3) ray-intersection-driven cross-view densification coupled with view-difference-aware adaptive densification. While preserving real-time rendering performance, our approach significantly improves both novel-view image fidelity and 3D geometric accuracy. Extensive experiments demonstrate strong generalization across diverse scenes and mainstream 3DGS variants, outperforming existing single-view methods in both qualitative and quantitative evaluations.
This work addresses single-image light field (LF) synthesis without multi-view inputs or specialized hardware. We propose an inverse image rendering framework that inverts a single RGB image into a set of source rays emitted from pixel locations, models ray propagation via a neural rendering pipeline, and explicitly captures geometric and semantic correlations among rays using cross-attention. An iterative ray expansion strategy jointly refines the source ray set while enforcing occlusion consistency. Crucially, our method avoids explicit 3D reconstruction and requires no scene-specific priors or fine-tuning, enabling strong cross-domain generalization. Evaluated on multiple real-world LF datasets, it significantly outperforms state-of-the-art approaches, enabling high-fidelity novel view synthesis, digital refocusing, and shallow-depth-of-field effects. To our knowledge, this is the first end-to-end, generalizable solution for single-image LF generation.
This work addresses the limitations of conventional monocular view synthesis methods, which are constrained by the pinhole camera model, by introducing a unified framework applicable to arbitrary imaging systems—including perspective, fisheye, and panoramic cameras. The approach operates within a consistent omnidirectional latent space, achieving continuous cross-camera view synthesis through joint implicit alignment of features and Gaussian primitives, coupled with a ray-based universal representation. Inspired by UniK3D, the model employs an encoder to extract 2D semantic and 3D geometric features, which are jointly decoded into a cloud of Gaussian primitives arranged along rays and radial distances. Evaluated on a new benchmark encompassing diverse imaging systems, the method significantly outperforms existing approaches, demonstrating strong effectiveness and generalization capability in generic monocular rendering tasks.
In real-world scenarios, uneven and insufficient human-captured viewpoints degrade novel-view synthesis quality. To address this, we propose a multi-scale contextualized visual guidance method. Our approach integrates semantic segmentation with vision-language models to identify key objects and prioritize them, constructs spherical proxy regions to encode viewpoint-dependent appearance details, and provides real-time hierarchical visual instructions to guide users in acquiring dense, spatially uniform image collections. Compared to conventional sampling strategies, our method significantly improves viewpoint coverage density and spatial uniformity in practical settings, thereby enhancing the visual fidelity of novel-view synthesis—particularly for emerging rendering techniques such as 3D Gaussian splatting. The core contribution lies in the organic unification of semantic understanding, geometric proxy modeling, and human-in-the-loop guidance, establishing an intelligent scene scanning paradigm tailored for high-fidelity 3D reconstruction.
Existing 3D surface reconstruction methods suffer from inherent limitations between volumetric and purely surface-based representations: the former often introduces redundancy and accumulative errors, while the latter, constrained by a single-layer receptive field, struggles to recover complex geometry. This work proposes a differentiable mesh softening strategy that explicitly extends surfaces into a multi-layer translucent volumetric representation with a controllable 3D receptive field. By integrating splatting-based volume rendering with topological optimization, our approach enhances representational capacity and numerical stability while preserving the advantages of surface parameterization. The method enables end-to-end high-quality reconstruction, significantly improving geometric accuracy and mesh quality in approximately 20 minutes, and effectively recovers intricate surface details.
This work addresses the challenge of high-quality novel view synthesis from sparse-view images captured under unconstrained real-world conditions by proposing an innovative framework based on 3D Gaussian splatting. The method effectively handles unconstrained sparse inputs through three key components: reference-image-guided view optimization, transient-mask-guided pseudo-view generation using a diffusion model, and a sparsity-aware Gaussian replication mechanism. These components collectively enhance geometric and appearance modeling in sparsely observed regions. Evaluated on public benchmarks, the proposed approach significantly outperforms existing state-of-the-art methods, achieving a 17.2% improvement in PSNR, a 10.8% gain in SSIM, and a 4.0% reduction in LPIPS, thereby enabling high-fidelity 3D rendering from limited input views.